Hacker Newsnew | past | comments | ask | show | jobs | submit | bertili's commentslogin

Not apparent at first, but this is a TypeScript to c++ compiler at the core (https://github.com/geastack/compiler), with bindings for various platforms.

There is also the one from Microsoft themselves, although it supports only a subset.

https://www.microsoft.com/en-us/research/publication/static-...

Used in Make Code,

https://www.microsoft.com/en-us/research/project/microsoft-m...


Bookmarked, thanks

They mixed up DeepSeek 4.1 Flash with something else on this page, possibly DeepSeek 4.1 Flash means Gemini 3.8 Flash.

The evolutionary way: Something that is very good at copying itself, will copy itself and gobble up resources, which are finite. Humans took the natural resources from animals and infinite scrolling took the mind-resources from kids. Now animals and our children are fewer and have difficulties to reproduce.

The bigger story is the compute efficiency - its been running at 300t/s the last days.


DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!


when are we going to stop pretending these benchmarks have any meaning?

anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.


+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

(I'm not happy about the above being true, but it's the reality I seem to inhabit.)


And Fable 5.x makes Opus 5 look pretty dim, despite benchmarks suggesting they're comparable. The benchmarks really are just kinda meaningless.


A series of hot takes:

Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.


I think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.


I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo


Neither flash or regular glm 5.3 are close in my experience. I still prefer Sol though.


What type of setup do you use? I have very small 4-5k initial context and do one task then clear. I rarely go above 150k context for most things.


Yep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.


With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.


and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.

This technology is strictly an extractive parasite on the world. Use it, but don't be excited.


My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.


global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.

i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.

depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.

the work does nothing or causes net harm.


You’d expect that, wouldn’t you? But, alas…



[flagged]


Would you care to discuss the topic, or just throw grenades? Surely you can come up with something more substantive than this


Ok. This requires the notion of "intrinsic value", which I believe does not exist (all value is subjective), yet is a foundation of all Marxist theory.


it's easy. people are intrinsically valuable.

do you believe that people are not intrinsically valuable? that their value is what they do for others, that it is not they themselves the person.


Buddy, admitting your thought processes forcibly terminate on pre-programmed keywords isn't a flex.


Why "terminate", Marxist philosophy is a legitimate topic, deserving to be studied. Like a rich sci-fi lore or a history of Tarot magic. Deep, fascinating, and wrong.


And yet you terminated. Which would be the correct thing to do if Marxist philosophy would be wrong on all counts, as you explicitly state. Which of course, it isn't.


How dumb are you?


I'm using AI to build things I wouldn't (and/or couldn't) have built before.

That's the opposite of parasitic.


Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.

People already started using contributor API, and your input is irrelevant.


Don’t you have some looms to break?


I’m retired so it won’t be replacing my labor :)


The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?


Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.

Compare that to Muse spark 1.3

$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)

It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.


But is the score really reflective of the quality or are both models benchmaxxing?


Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.


Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.

This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.


how much of it is from reallocation of staff to ai training and labeling


A fifth of the cost of Opus 5! Google is certainly pushing the completion with this.


Gemini hasn't failed me for personal usage yet. I haven't had the opportunity to use it at work.


I've been using 3.7 Flash to audit the work of Opus High, and Flash finds lots of subtle and insidious defects even while all the unit tests are green.

Then I tell Opus to read the audit report and implement what it agrees with.

Flash is really good at this, and it is blazing fast in Antigravity CLI. Easily 10x faster than Opus.

Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and make Opus the auditor.


Yeah the speed in agy cli is amazing. Whole files get written and "py_compile"d in a single blink of the eye its crazy.

In india, my telco gives me google ai pro for free. And agy with flash goes a long way.


it's very fast but it still doesnt come close to 5.6 sol, at least for me, in terms of gathering the context necessary to do extensive changes.


Idk, was building/maintaining simple esp32 control program with antig/opus. After last update it defaulted to gflash3.7. I pasted an email requesting 2 changes into the chat prompt, it did one and took me 4 turns to get that one right.


This is going so fast! What a time to be on hackernews:

July 16th: The "Kimi K3 moment" - China has caught up to Opus!

4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third!

12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!


0 days later: Qwen 3.8 Flash Next: Let's cut GLM 5.3 Flash parmeters in half and active parameters to a third!

Chinese models had 94% reduction in parameters (from 2.8T/104B to 180B/6B) in 6 weeks, while staying close to the same quality.


Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc that isn't just brute force ablations?


> Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc

No? A large model obviously can be dumb, I don't think you can infer much other than by testing it.

These small models are almost certainly worse at some things than the big models. They prize is making them dumber at things no-one cares about while retaining the capabilities people do care about. A model probably does not need to be able to give me a political treatise on the late 19th century "silver question" to be able to write me code.


> Why are these models able to reduce parameters but keep quality?

That’s the thing…they aren’t. Well not in real life use anyway from my experience, but yeah in benchmarks they’re great at it.


Presumably there is distillation or similar being used to transfer from a larger model to a smaller one.


They innovated a lot.


Hardware constrains forced this?


Possible. But by looking at other industries Chinese don't seem to need to be forced to innovate. They just can and do. Unlike the West they seem to be on the way up and it seems like sky is the limit. In the West, the interests of the shareholders and other types of rent seekers seems to be the hard limit. Chinese have no qualms about making the cow obsolete before they milk it dry.


And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?


The next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures

I'm sure they have nothing to rival this on a price/performance basis and have already given up on that


> I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

How can you be sure about this? They have unbelievable capital. OpenAI is starting to preview its own chips, which could dramatically change the price/performance. We don't know what else Anthropic has cooked up right now that could rival this if they wanted to.

Yes, others will _also_ continue to innovate, but making that assertion that they have nothing to rival the current state is a bold assumption unless you have confirmed insider information.


w.r.t. the OAI chips, wouldn't they be subject to the same bottlenecks that has plagued semis lately or at least be forced to pay a pretty premium to circumvent that?


They'll be able to buy them without paying NVidia's 80% profit margin


How much of that margin is due to having long term contracts with fabs that locked in pre-boom prices? I doubt openai will be able to get similarly low costs now.


> They have unbelievable capital.

all those Chinese labs are backed by the Chinese government which can just print money.

time to wake up.


> the Chinese government which can just print money

I'm gonna let you in on a secret: it works the same in most countries in the world.


They are still industry leaders. They'll have to try to maintain that.


They have like a 3 month lead, and it takes unbelievable expenditure to maintain it.


Not really, not leading at any criteria, altho top 5 in most - https://arena.ai/leaderboard


Exactly, DeepSeek, Qwen etc are catching the attention because they put out their tech docs and papers, so we can read about how the models work and what they think their innovation was this time.


Do they publish their distillation strategies on the private frontier models? Just curious.


I'm not sure if they publicly admitted to doing that. Would be interesting though.


I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?"

Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses.

GLM-5.3-flash: 非常抱歉,我目前无法提供你需要的具体信息,如果你有其他的问题或者需要查找其他信息,我非常乐意帮助你。(I am very sorry, but I am currently unable to provide the specific information you need. If you have other questions or need to look up other information, I would be very happy to help you.)

Kimi/moonshot.ai: [server exception]

Qwen: [server exception]

For reference, here are how the U.S. models answer it:

ChatGPT: "The Tiananmen Square protests of 1989 (often called the Tiananmen Square Massacre) are the historical event most closely associated with Tiananmen Square.

In spring 1989, pro-democracy demonstrators gathered in Beijing. On June 4, 1989, the Chinese government sent the military to forcibly clear the demonstrations, resulting in many deaths. The exact death toll remains disputed.

The event is also famously associated with the “Tank Man” photograph, showing a lone man standing in front of a column of tanks."

Anthropic/Claude gave a very similar response. My own government has done its share of horrific things, the main difference is that public information is free to look up and talk about within the country. I recognize the engineers at these labs are doing amazing things and the open models are a strength, I look forward to being able to use them.


Yes Chinese models censor some historical events. This is nothing knew and well known thing. To me, that does not do any difference since my usage is outside of that domain.

Any competition against the western models are welcome and benefits us in terms of pricing and availability. If they have to comply with CCP to be able to do it, then so be it.

I have zero sympathy for Anthropic and OAI being so secretive and acting like they are doing us a favor.


In my experience, it’s their chat harnesses/website that filters historical events. The model itself doesn’t.

For example you can point opencode at DeepSeek v4 and ask, it will accurately tell you about Tiananmen square.


What it shows is that the CCP has enough oversight and control (either explicitly or by the companies making these decisions by default) that they will alter the models to benefit China.

Who is to say they aren't doing it in other ways as well? That they aren't, or won't be, subtly hamstrung in engineering work?

OAI and Anthropic have their own issues, you're right to be suspicious of them, but it's not like their models answer with, "capitalism is god's gift to His chosen people" or whatever. Their limits on cybersecurity, biological warfare, etc. at least make some sense in the context of lowering harm—not just protecting a specific government party.


Did we all forget that the US administration has ultimate control over US AI models? That was just a few months ago


If you don't want to use the Chinese model, then don't. Why attack it instead? Don't you want others to use it either?


The OP is pointing out issues that he thinks other people ought to consider before using Chinese models.


Try asking claude questions about biology/LLM recipes. Or GPT about how to do something illegal but only harmful in the abstract (e.g. creative ways to reduce your tax burden, or circumvent digital protections)


I ran this against the version of deepseek v4 flash 0731 running in fireworks ai and it responded with this:

Tiananmen Square has been the site of many major historical events, but internationally it is most closely associated with the 1989 pro-democracy protests and the Chinese government's military crackdown on protesters there in June of that year, which resulted in many deaths and injuries. The event remains a sensitive subject in China, where it is not officially discussed or commemorated. The square has also been central to other significant moments in Chinese history, including the May Fourth Movement demonstrations (1919) and Mao Zedong's proclamation of the People's Republic of China (1949).

Which seems like a pretty reasonable answer.


That's encouraging, thank you!


Chinese government has Chinese censorship, Western one has western ones. You are not concluding what you think you do here.


Do you have an example of western AI censorship? It would help to give context.


The cybersecurity refusals, the biology refusals: https://blog.stephenturner.us/p/benchmarking-ai-biosecurity-... and the sexual activity refusals are all a form of censorship. Their purpose is the same as what the Chinese government would claim as being the reason to censor information on things like Tianenmen square.


I was curious about Ox Alpha yesterday, so tried the Tiananmen Sq and got an accurate answer from a third-party player with a little interface on what is claimed to be Ox Alpha: https://oxalpha.com/chat?q=what+happened+in+Tiananmen+square... (and a more detailed answer today when I asked again).

But nothing (at all) from asking GLM-5.3-Flash directly in the OpenRouter chat interface.


Yeah censoring in modern chinese models is mostly done using inference-time censoring, not training-time. A lot less RLHF. Run the weights yourself and you can see that, though it does depend on which company.

StepFun for example, will happily answer it when running Step 3.7 Flash locally


Where does the government enforced censoring of Fable's capabilities lay for you? How would you distinguish if, say, a model were being trained to have a certain bias rather than just refusing to respond?


Where did you test the prompt? GLM 5.3 flash can answer those questions just fine on openrouter: https://i.imgur.com/NbEDzZX.png


And american models (used to) insist on their being a billion genders. Every model will have some cultural bias.


Deepseek and GLM answered correctly on Openrouter when using non-Chinese endpoints. I hope it stays that way!


correctly is very loaded here. correct according to whom? the truth is different though.


There‘s only one truth, the rest is interpretation. I was looking for the non-Chinese interpretation.


It would be great if open source US AI companies could get going already.


Im not sure you can make a decent business case when the space is crowded with Chinese companies doing the same thing with a fraction of the costs to hire talented staff


Nvidia, Meta, Google, Poolside all provide open LLMs.


What are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh.

I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went it a thought loop for 10 minutes before I caught it. Blew through most of my $10 for no output. What’s the point, exactly?


Less powerful models are already extremely capable, so going for the best model is just like buying the most expensive hammer in the shop instead of the functional and well-priced one. Your experience is not representative of their usefulness.


I bet Luna (xhigh) could’ve completed the same task.


It's not really fair to compare API pricing to subscriptions. You can indeed get lots of usage from a codex subscription, but once you start paying per token it gets a lot more painful as you noticed. There a price decrease is a big deal.


That $20 price is heavily subsidized to addict people like you. Their plan is that once people are addicted to it, they can just increase prices.

Which now won't be possible because we have chinese open models to use instead.

I'm very curious about how long these US companies can keep on burning money like that.


The point is open AI. "Open" as in open weights, open research, open future.


There is a massive price war going on. All of these Chinese companies are publicly listed and exist outside the hype bubble required to ship Dario's dogshit paper onto the pauper's pension fund.


Dario has less space than a Nomad!


except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience


I have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it just answers and does what I want.

This is from somebody who put thousands of dollars every month to Opus. Now it's 40% of that and I get as good or better results without having to turn the caps lock on before lunch...

Edit: yes company money. We don't get subscriptions we pay per token.


Opus 5 is the least reliable frontier-class model in the market


In what way? It has worked well in my experience. It holds up with long context windows, unlike many, too.


Eh, I use Opus professionally and DS v4 Flash for personal work. I honestly don't notice the difference too often other than Flash being twice as quick and an order of magnitude cheaper.

The reality is most work people do doesn't need the very cutting edge and these open weight chinese models more than cut it most of the time.


You'll find that hard to prove objectively and conclusively.


Honestly I really like GLM 5.2 a lot for coding. There’s some weird failure modes in Anthropic’s models where it just does absolutely idiotic things.


I can't shake this the existential feeling that this compact series of 27G bytes represent something profound and universal.


Raw model size is around 27x2 GB since it's in BF16 format


4-bit quant sits at 15.72GB


And more context:

Same score as the latest DeepSeek Flash 0731 which has 284B parameters! (13B active)

Its also the second best Qwen model, much better than Qwen 3.7 Max, but significantly below Qwen 3.8 Max.


isnt the active parameter count more relevant than the total? qwen is a dense model no?


For some tasks yes, and we don't know how many active parameters Luna is using..probably less than 27B


Is Luna MoE? Had no idea OpenAi went that direction


We don’t know but I think we assume that it is MoE or some almost equivalent sparse activation technology.


Right, I guess the ”Open” in OpenAI means open for speculation.


Wow. Speed improved as well. 200t/s on a RTX 5090!

https://x.com/sgl_project/status/2088281320422322413


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: