This is awesome. I am trying to build a full scale ASR system within 20-25MB. Now that we have Claude code to run experiments, I have started running some experiments. Promising results so far. First realization is that you can capture the nuances of speech in just 3300 embedding vectors(786d). This sequence can be decoded with a small CTC system to get text. Next experiments are on reducing the 768 dimension space into a 64D space. Thats also show some promising results. Hooking up my system so that the agent blogs the results everyday[1]. So my research "claw" setup does the experiments and posts results which I check in the morning and adjust the experiment direction as needed. Its not fully automated yet, but almost there.
I think Google's Conformer paper is SOTA at the <30M model size, where I think they put an incredible amount of flops into a 10M param model to reach around 2% lsc clean (the whole model and RNN decoder were trained domain specific to librispeech here).
I think my small Talon models are next, around 3% lsc clean at ~28M (greedy CTC decoding, no external encoder, no LM, not trained in a domain specific way). I reached around 6.5% at 10M.
I've been working on some new baselines I want to release soon as public artifacts. This article is inspiring me to try pushing the param size down a bit. I suspect we can do large vocabulary end to end in the <5M range.
You can get very close to a working version by reverse engineering the javascript with Claude. I got a good version almost an year back. With Opus it might do a better job.
Most of the problems happen because we want to simulate human conversations. While thats a good goal to have, another approach is to let the user know clearly they are talking to a bot. You will be surprised at how accomodating users can be when they know they are talking to a bot and want their queries resolved.
Lets break this down. There is very little in newness in what Anthropic announced. Claude had skills for a long time. They have added one more layer of abstraction and called it plugins. This mainly comes with a set of integrations.
Thats the pitch.
But, what are Claude plugins?
Plugins=Commands+Skills+Integrations.
Commands are specific to Claude code. But commands and skills are nothing but prompts at their basest level.
So what is the main differentiator?
Integrations.
But what are you integrating with?
SaaS companies.
And what is the stock market doing?
Dumping SaaS stocks.
How do they think Claude cowork will work without the integrations. Without the system of records.
If anything, these SaaS products have become more important. If I was a trading guy, I would go to the github of claude plugins, see the default integrations and buy the stock of those companies.
Claude cowork and the SaaSpocalypse case makes no sense.
What are Claude plugins?
Plugins=Commands+Skills+Integrations.
Commands and skills=prompts
So differentiator?
Integrations.
Integrating with?
SaaS companies.
And what is the stock market doing?
Dumping SaaS stocks.
Organizations adapting AI is the biggest problem that businesses are facing right now. Even in Ozonetel I face this problem day in and day out. The employees who really use AI to its full potential are minuscule. I can count on my fingertips. We need to overcome this in the right way or we will face the same problems we faced during industrial revolution.
Got pissed off with too many ARR manipulations and AI startups announcing revenue numbers manipulations(best day * 365 as ARR etc). These startups are not only messing up the AI ecosystem with the non standard numbers, they are also messing up the SaaS ecosystem by co opting the SaaS metrics. Now SaaS startups are supposed to show the same scale though the AI startups have also not achieved that scale. So here is what I propose, the AI startups should use their own new vocabulary. Since everyone is vibe coding, I suggest VRR, Vibe Revenue Run-rate :)
I have even provided a formula(all scientific and all) and also provided a checklist for the VCs.
reply