“when committed as part of a widespread or systematic attack directed against any civilian population” which I don’t think this really clears, as objectionable as his comments may be
DHH hasn’t committed a crime against humanity. He just said he would not consider one entirely bad if it made him feel safer when walking around his house.
I don’t recall him being specific about which Roma are wolves and which aren’t. Were he talking only about “criminal” migrants, he’d make a point about that distinction.
I don't think he's been literal about ethnic cleansing, even though I strongly disagree with his political views on the matter.
Anyway, he also supports open weights models, which IMO is a good thing.
So maybe separate the artist from the art? He's not running for elections, that's for sure, so as a citizen he can have the political views he likes, doesn't mean the software is bad because of them.
As I've said before, I strongly disagree, but what I read is that he calls for the removal of those behaving in a way that (in his opinion) make the city of Copenhagen less safe, or at least that change the perceived security of the people living there.
Gypsies create issues almost everywhere in Europe, but it's not based on ethnicity, most of the ethnic Roma are fully integrated, some of them are not, in Italy many Sinte families are linked to mob and criminality in general.
Do they deserve special treatment because of their ethnicity?
And what if they are not regular citizens?
It's an open problem, DHH's solutions it's absolutely not my solution, but there's a still a difference between what he said and calling for ethnic cleansing.
No one will reply to you because it requires thinking through this problem with the same nuance they accuse DHH of not appreciating. I'm right there with you.
Remigration is a far-right concept referring to the ethnic cleansing via
mass deportation of non-white minority populations, especially immigrants
and sometimes including native-born citizens, to their place of racial ancestry.
I'm not interested in having "nuanced discussions" about whether the Roma people have some kind of Nuisance Gene that makes discussion of extermination valid
If you'd like to make an argument that isn't just "ethnic cleansing is okay sometimes if the ethnicity is REALLY annoying" go ahead
I specifically use Android because of this, among other requirements, that Apple imposes on software development on their platform.
I do not own general purpose computers that I am not allowed to develop software for without permission. I have always avoided consoles for that reason as well (Steam Machine and other similar platforms would be fine, but I've been avoiding consoles for long enough that it's not something I really look for any more).
I was a major Apple fanboy up until the iPhone. Left the ecosystem after the iPhone and macOS started moving in that direction as well.
I'm going to miss having a smartphone that I can use with my banking and EV apps, but probably for the best to get out of Google's ecosystem. Hoping that GrapheneOS will still allow me to use some of the apps that I like.
I switched to an iPhone last year, mostly because I was getting sick of Samsung’s shit, and google doesn’t sell their pixel phones locally (and I don’t like Oppo).
But the other reason is that most of Android’s openness is quickly disappearing, so my main argument against iPhone is gone. And on the topic of privacy, I actually trust Apple way more than I trust the company that makes most of their revenue via advertising.
Exactly this. Once the writing was on the wall that deGoogled / FOSS Android was on borrowed time (at best), a ton of the argument against iOS dissolved overnight. We can argue that Apple's anti-repair policies and anti-user-customization policies are their own evils, sure, but at least my phone works for its core tasks, and works phenomenally well at that. My deGoogled Android phones were mostly hobbled together piles of "it sometimes works, as long as I don't look at it too funny". This tradeoff was fine for fully owning my data and being able to install absolutely anything I wanted on my phone. With arbitrary APK installation disappearing, and unlocked bootloaders to install a less hostile fork of Android becoming such a rarity these days, I may as well at least not fear looking at my phone the wrong way when the moon is in alignment with the wrong star.
Source: about 12 years of Android usage (11 of those on unlocked bootloaders and custom ROMs, ~6? of those deGoogled) -> iPhone 16e.
For the first half, custom ROMs were a massive boost to the experience, but I never really bothered after 2018; You had to install so much of google's stuff to get anywhere near proper experience that it just wasn't worth it, and even then many apps were starting to throw a fit when they detected a rooted device.
And at the same time, the stock android experience from most vendors was now good enough. I didn't even check custom ROM compatibly for my last Android two purchases (both Samsung). But it was nice to have the option of rooting and maybe finding/making a custom ROM Samsung removed that in their last update to One UI :(
For me it's the data-hoovering. I don't want Samsung (or Google, or Xiaomi, or whoever) background services phoning home my location every few seconds, or being able to uniquely track my device across apps, etc.
Once my ability to run a "background spyware-less" Android device went away, so did my willingness to put up with the negative sides of the usability tradeoff.
Apple probably sends some telemetry home, but at least their business model doesn't depend on knowing the exact Pantone color of the food I ate for dinner to market a matching wall paint to me.
Edit to add: oh, yeah, you mentioning 3P apps complaining if they detected root or MicroG reminds me of many a horrible night of debugging. Ugh. What a mess.
> You had to install so much of google's stuff to get anywhere near proper experience that it just wasn't worth it
What was it? Personally, the only thing I miss from stock Android is the Google Wallet/Pay/Wallet. (I do use microG to get push notifications and embedded maps working.)
Apple already have a huge chunk of the ad pie and are growing all the time. Dont think Apple are keeping you safe from ads and all the profiling of your behaviour that involves.
Apple are still far less reliant on advertising than google; They take far stricter stances against apps tracking users than google; And they actually let me an Adblock extension (or any other extension) in Safari.
Apple is not some magic solution to the question of on-device privacy. But compared to Android, the improvement is night and day.
Android has been a brutal disappointment on this front since the day it launched. Sure, you could hack various devices and spend too much time on xda-developers and get custom ROMs, but if you compare that to the openness of even a stock Windows PC it's an absolute joke. Android was supposed to be the open Linux phone, when I bought my HTC G1 full of hope; turned into an inferior iPhone wannabe with worse performance and a low quality walled garden that keeps most people in without keeping the trash out.
Ironically, LLMs are pretty good at finding stuff like that from vague descriptions. With that, I think the parent commenter meant Art Graesser's AutoTutor work based on the terminology used (hint/pump).
TLDR: Actual human tutoring sessions were recorded and analyzed and Socratic questioning was barely used at all. Instead the following pattern was observed:
Pump — "Uh huh?" "What else?" Costs nothing, so try it first.
Hint — points at the region of the answer. "What about the pumpkin's motion sideways?"
Prompt — fishes for one specific word, with the sentence frame supplied. "The pumpkin keeps moving forward at the same ___?"
Assertion — just says it. "It keeps the runner's horizontal velocity."
Yes, you're on the money. And this is a perfect example of how people should use llms. Use them to deepen your own research and understanding rather than use them to glaze and implement the first naive idea that comes into your head (or the first idea they suggest as a response to a vague prompt, which typically is substantially worse than if you get them into the right vector space using jargon)
They do maintain the Transformers library which is pretty much the core library for how you interact with LLM models in the open source world. So while they weren't using a model they've trained, they were a part of making just about all of the open models (maybe excluding OpenAI and Google's, I wouldn't be surprised if they have their own frameworks that predate the Transformers library).
3.6 Flash scores exactly the same as 3.5 Flash on the Artificial Analysis index. Better on some tasks, worse on others. Mostly within what I'd consider the noise window. Looks pretty much indistinguishable from 3.5 Flash, at least on these benchmarks: https://artificialanalysis.ai/models/gemini-3-6-flash
> I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now
10-15 years? The current rate is closer to 10-15 months.
15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.
Today, you can easily run Qwen 3.6 27B on a variety of consumer hardware. It scores 37 on that index.
I've run all of these models on my laptop (Strix Halo, 128 GiB of unified RAM); the bigger ones, like MiniMax M2.7 and DeepSeek V4 Flash, need to be done at fairly aggressive quants that will certainly lose some performance and not quite hit the performance of the unquantized models. But still, it's definitely the case that you can run models that are competitive with the frontier models of 10-15 months ago on consumer laptops.
Heck, just announced though the weights haven't yet been released for independent confirmation is MiniCPM5-2B, a 2 billion parameter (small enough to run on your phone) model, that according to their benchmarks has performance competitive with GPT-4o, a frontier class model from 2024.
So that's around 1 year for frontier to consumer device class, 2 years from frontier to phone.
Now, this kind of rate won't necessarily keep up; it's possible that local models will hit a performance ceiling before frontier models do. There's only so much information you can cram into a certain number of bytes, and the AI boom is causing hardware prices to skyrocket so keeping consumer hardware from advancing quite as fast as it had been.
No. We need objectively around 192 to 512gb of very fast memory to be able to run really useful models. I don't see local hardware with these specs coming in 1 to 2 years. There are a big number of initiatives currently taking place to increase ram output. But it will take another 3 years minimum to close the current supply issues. China is fast pacing forward to have its own chip baking factories with small enough nano scales to have fast chips. Will also take a few years.
> 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.
There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors
12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore.
Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.
Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).
Pretending progress hasn't been mindboggling is insane.
Yes, if they are made with new technology and the quality is good enough they certainly are.
AI generated video memes were NOT a thing a year ago. Yes, the AI video generation in itself was a meme (Will Smith eating spaghetti), but now there are tons of convincingly good AI video memes about things like sports events generated by average Joes. It is proof of the capability and accessibility of the technology even if the use of it is mundane.
Everybody and their dog playing snake on their Nokia 3310 was similarly mundane, but also a sign of the end of the era of Gameboys and the beginning of (normie) mobile gaming.
> 10-15 years? The current rate is closer to 10-15 months.
The leaps between models have gotten smaller and smaller. 2023-2024 models were rocketing up in quality. 2024-2025 I’d say was pretty impressive too. But 2025-2026? Very easy to feel the slowing pace of improvement. I agree 10-15 years is overly conservative but 10-15mo is far too bullish.
The speed of model releases, in my view, is actually getting faster and faster. There were nearly nine months between GPT-3.5 and GPT-4. And now in just over one month, major models already included Claude Fable 5, Claude Sonnet 5, the GPT-5.6 series, Kimi K3, GLM 5.2, Qwen 3.8 Max, Grok 4.5... and the official DeepSeek V4 release is coming soon.
Both Codex and Claude Code have it, but they work slightly differently.
Claude Code uses Haiku to read through the transcript and decide if the goal has been completed. If not, Haiku injects a prompt back to the main model to indicate what still needs to be done.
In Codex, instead it's a tool available to the main model, plus some part of the surrounding harness that will re-prompt it if the tool calls haven't yet indicated that the goal is complete.
The issue that they are trying to solve is that sometimes models will stop before they have actually fully completed whatever task they were given; attention isn't perfect, and someitmes they'll complete part of it but not the whole task. Rather than making the user come back and re-prompt to keep going, they add a way to automatically do a bit more nudging to try to get the model to finish the task.
> Claude Code uses Haiku to read through the transcript and decide if the goal has been completed.
feels kinda odd to use a less capable model to determine if the goal is fully complete. Especially if the user is expecting /goal to thoroughly complete the task. A less capable model would be more likely to misclassify `isComplete?`
I've tried doing a loop of rending the SVG and then tweaking based on that, with local models (so, not nearly as strong). It wasn't very successful; it would mostly report that the image looked great and didn't need any tweaks. Maybe I should try it again, there have been some newer models since I first tried it. And yeah, maybe worth trying with bigger models. But I have found that models aren't necessarily the best at visual reasoning and review, even with a vision loop. Their lack of visual reasoning is part of why they still have trouble with things like ARC-AGI-3.
I've found much better luck giving it an audit check-list, including some steers like: are there any visual glitches or SVG bugs, are the colours consistent, etc.
You can always ask them to draw something else, as a way to avoid any possible pelican related data contamination; given how popular the pelican test is, I'm sure there's some pelican SVG drawing in the training sets of at least some of these models by now. For instance, you could ask for an SVG drawing of a cyborg bear riding a rocket powered unicycle.
It's a silly fun little benchmark, and because Simon's been doing it for so long, you have a lot of examples over the years to compare. But you can always come up with and run your own test with other drawings.