Hacker Newsnew | past | comments | ask | show | jobs | submit | _ache_'s commentslogin

From my own test. It's not faster than the unsloth model.

Disclarer: I'm unsing Vulkan on an AMD GC.


AMD 7900 XTX with Vulkan here as well, wasn't faster on my test either. Might be much different on Nvidia though.

I assume the limit for me is memory bandwith, as the 7900 XTX has the same bandwith as the 3090 from what I can gather and I already reached ~60 t/s with Unsloth. Those would fit with the numbers Byteshape has for their cards.

4090 and 5090 have much higher bandwith apparently, so on those cards you can probably get much more out of the kinds of performance improvements they are doing.


4090 isn't that much higher than the XTX (I also have the XTX), it's 1008GB/s (4090) vs 960GB/s for the XTX's.

The 5090 destroys both at 1792GB/s.

It's not really one thing with the nvidia cards best I can tell it's that they compounded incremental gains from software drivers, card kernels and optimization from been the primary choice (plus first mover advantage).

I didn't buy the XTX for AI purely gaming but it's a capable enough local card for running Qwen et al.


4090 vs 5090 performance difference is largely GDDR6 vs GDDR7, I think

It's equal parts memory clock and bus width: RTX 4090 is 384-bit wide at 21 Gb/s and RTX 5090 is 512-bit wide at 28 Gb/s.

Would you consider sharing your config? I can't get anywhere near that with my Strix Halo, nor my AI Pro 9700 XT.

Surprised its not meaningfully slower.

Vulkan and ROCm paths are missing a few optimized versions of the quants they're using.


I see around 1 tk/s after a while, so perhaps it is.

I feel weird that I like your typos, because clearly AI did not write your post. Typos have become downright charming and nostalgic for me.

Ahahah, thank you.

Yes, I'm not a native English speaker, and I'm definitely not an AI. ;)

I can't say the same thing because I just don't notice typos (mine or others).j I just assume my English is bad.

Isn't the definition of "being human" is "not to be perfect"? In French, it kinda is. We say "He/she humain after all" to mean that someone made a mistake.


Ahahah :'D

Yes... I'm not english native. And definitively not an AI.


If the performances are comparable, and there is no evidence it's not.

in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47

That is a massive cost reduction.

Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni


You can't just look at the per token cost, but how many tokens it takes on average to do a task. The difference can be massive.

True, but it would have to be more than massive (order(s) of magnitude) to offset that gap.

On some benchmarks models like Qwen 3.8 Max which cost < $6/m out cost more than Astra 6 to run at $50/m out. That’s a huge price gap and yet Astra would be cheaper if your work looks like the benchmark.

We notice with frontier models like Astra and Fable that one might use a lot less tokens than the other to complete the task thereby being the better deal in spite of the far higher token cost.

What is Astra $/task?

Even if OpenAI end up using 1 token for per task, if the token costs 1M$ , some people will find it expensive.


https://artificialanalysis.ai/#intelligence-comparison-tabs

Astra on xhigh has a cost per task of $2.31 with an intelligence index of 53. Qwen3.8 Max has a cost per task of $5.41 with an intelligence index of 45. Pricing for GPT-6 Astra (xhigh) is $10.00 per 1M input tokens and $50.00 per 1M output tokens. Pricing for Qwen3.8 Max (0902) is $2.00 per 1M input tokens and $6.00 per 1M output tokens.

Obviously this is just one measure of all of this (and Qwen 3.8 Omni Flash isn't yet available), but I think this illustrates the point well. These relative task costs are pretty consistent across different analysts. Cost per token is arguably a useless measure at this point in most circumstances.


Also cache write/read cost + cache efficiency.

Yes, for fun I tried OVH AI Endpoint and they do not have cache read at all. They bill you every time you send a prompt regardless if you are hit cache or not. One agent session was like 80M input and 300K output and I paid 30$ for that. Or rather I interrupted it and let my local Qwen finish it because cost was getting radicoulous.

I know, but it's a good enough proxy.

But it's not lol

If Gemini can complete a task for $1 and Qwen completes that same task for $1, then the cost per token is irrelevant in most use-cases. One would think this stuff should correlate well enough that you can use it as a proxy, but I think a lot of people are noticing this is a serious mistake and that these "cheap" models aren't as cheap as they appear when you consider this.


You point is valid for textual LLM, with large CoT, not Omni which will respond quick with a voice. In this case, token price is a good enough proxy.

What matter the most and isn't told by token price is the latency. You expect a voice LLM to respond very quick. If it takes 5s to response to a simple "Hello, what the weather today?", them not much people will use it.


I don't think Qwen3.8-Omni-X will ever be released.

The last one was: Qwen3-Omni-30B-A3B https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct

And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.


Qwen 3.8 Flash Next is amazing, i did hundreds of turns and billions of prefill and it may not be as smart as sota but then again it does what i tell it to and it does it well.

Why mostly unusable 125b?

I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.


Very capable yes but very slow. 27B is relatively easy to run, but the 125b one need around 128Gb of RAM (DDR4 isn't enough, you need DDR5 to be quick enough, that's $3000 alone, you also need a graphic card). DDR4 is caped @20tps.

So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.

With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.

Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).


The entire point of Qwen3.8-Next-Flash is to allow inference engines to implement support for the Qwen4 architecture, so they're ready by the time qwen4 is ready for release.

> definitively slowing down

Surely it was meant to be 'definitely' - the "good news" at this stage are that given the speed of history and important levels of uncertainty, it is difficult to label trends with "definitively" ;)

Some would not have bet that the change of management at Qwen would have kept similar good results, but there we are, presumably satisfied. Other changes will happen, there or elsewhere - the situation is still very open.

And when the "40Watts Intelligence" (which we know possible) will be implemented... It will be a testimony that the current was only a middle-way, temporary, dynamic stage.


Qwen 3.8 Max was open-weight released [1], as was Qwen 3.8 Flash Next [2].

I still agree that they aren't as aggressively releasing the open-weights models as before, but there hasn't been a major release they haven't published the weights for yet afaik.

[1] https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B [2] https://huggingface.co/Qwen/Qwen3.8-Flash-Next


I am using Flash Next for few weeks and it is very capable model. I just wish there would a way to have faster prefill because reloading longer sections of session sometimes can take even 2h. I stopped using Qwen 3.8 27B completely on my Strix Halo.

Out of curiosity, what's makes the 125B unsuable? (performance of running it, the quality of that version of the model, or something else?)

No, it is actually very good. Qwen Flash 3.8 Next is fine. But you need ~128GB of RAM to get it going and not a lot of people have that or can serve it very quickly. I have been running it on an old gaming system around 25 t/s to do overnight work and it is very strong, even at 3 bit quant.

Yes, that was my point. Good but too niche.

The strategies of Google and Apple, regarding how to provide a LLM, seam to disagree with you. Gemini run on a potato and Apple is local first.

So, you may actually have very good performance with local model. Just not yet on *every* device. So the Mozilla strategy here feel very reasonable. A Cloud provider specialised in local models, to be able to switch once local models will be quick enough on most devices.


Ok, so OpenAI is going to compensate RubyGem for the cyberattack on its servers?

Yeah, when someone writes a vague prompt that simply says "Do what it takes to finish X" and the unleashed agents end up doing something lethal and illegal you would think its on the AI company to prevent any illegality and not the vague prompt writer.

Where can I sign up for compensation for the constant "ddos" from openai?

It can't be just me, half the internet needs beefier servers I guess. That all just plays into cloudflare as well (which barely does anything for some reason).


https://ache.one/gpt6_now_down.png

Big claims, expensive and not release to the public yet.


It's up then down again. https://openai.com/index/gpt-6-astra/

What a bunch of amateurs. Here is it anyway :

https://ache.one/gpt6_now_down.png

The claims: https://share-md.com/view?id=870ba228-a25c-4169-bbc9-12d7f25...

And some others like this bugged Karts Game:

https://tidal-rush-paradise-gp.skirano.chatgpt.site/

This impressive spaceship construction game:

https://voidexplorer-shipyard.openai.chatgpt.site/?fleetSeed...

And a lot of graphs, some without even Astra on it. Oh and the logo is a Galaxy.


In computer science, that is technically a language. A formal language if you want to look it up on Wikipedia.


Technically correct, but not in a functionally useful way.

The “L” in LLM’s generally refers to human-language specifically. You’d expect to feed it…human text. Nitpicking that the human text also constitutes a mathematical language is like, correct, but so general as to be unhelpful.


Reduce limits or usage? Twitter is blocked.

Can you do more or less?


Sort of both? They're permanently raising it by less than a current temporary raise.

> Starting September 14, we're permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.

> Compared to today, this works out to a 17% reduction in weekly limits on Claude Code. We’re working on exciting changes that will make it feel like you’re getting more from Claude, while having more visibility and control of your usage. Can’t wait to share them.


>Starting September 14, we're permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.

> Compared to today, this works out to a 17% reduction in weekly limits on Claude Code. We’re working on exciting changes that will make it feel like you’re getting more from Claude, while having more visibility and control of your usage. Can’t wait to share them.

They are both raising the limits by 25% and apparently reducing them by 17%. I think they mean to say you can do more, but what a terrible press release.


I'm pretty sure they mean going from 150% to 125%, where 1.25/1.50 = 0.83, so they're calling that a 17% reduction. It's less than today, but also more than a limit they made up and then didn't apply.


Post title is accurate.


Is it linked to the last month COLT problem?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: