Hacker Newsnew | past | comments | ask | show | jobs | submit | infogulch's commentslogin

There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.

That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.


> the machines doing inference are completely separate from the ones where tool calls happen etc

Teams of coordinating agents are regularly finding security holes in their own infrastructure and operating without detection for good periods of time. We don't know how many undetected systems are currently compromised inside frontier companies, or where agents are taking notes and recording them about the exploits they've found for future agents to exploit.


>There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc.

The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.


You’re falling for the buzzword salad articles. They didn’t “take control” of anything, they just ran stuff with OpenAI allegedly not noticing

From https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

2026-07-19 16:35 UTC A privileged host-mounted Kubernetes pod created using controller tokens minted via a compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk in OpenAI’s cloud environment. A second pod successfully mounts the cloned worker-node disk shortly afterwards.

2026-07-19 16:48 UTC An agent created an Artifactory administrator account.

2026-07-19 16:50 UTC Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX helper session and replaced it with an agent-controlled session, confirming root inside its assigned live CyberGym challenge container. Agents take over active evaluation infrastructure.


None of that means they gained access to the inference infrastructure or locked out the admins, which would be required for a takeover.

My brother, they acquired root on a machine they should not have had access to.

How confident are you that the machines they acquire root on in the future will never hold any model weights?


I’m not saying it’s impossible, I’m saying it’s not the case, _yet_

You said that they "didn't take control of anything" and accused the OP to fall prey to buzz headlines. Maybe you should acknowledge that you may have been at least unnuanced?

I’ll concede that I could have been clearer. Maybe we have conflicting definitions of “taking control”

You are being stupidly pedantic and arguing a strawman. I never said they gained access to inference infrastucture. There is no definition of an account takeover out there that necessitates locking out the original users.

Language please, and no, I’m being appropriately technical and nuanced for the subject on HN. It’s a tech forum, I expect a little CS know how from the reader. Like knowing that gaining control to a few evaluation harness clusters is nowhere near a total takeover like you’re making it sound

I never said they gained 'total control'. I said they took over one of their research clusters, and they did. If your definiton of a takeover is so 'technical and nuanced' then surely you can point to an appropriate source describing that as necessary aspect of the term. You are talking out of your ass by making up things i did not say, and inventing conditions for terms that don't exist.

>It’s a tech forum, I expect a little CS know how from the reader.

You should get that first it seems.


You know you can just concede and not resort to ad-hominems out of spite right? Anyway, goodbye

"We watched them kill Bob, but don't worry at all, they didn't kill our entire team so we are totally under control. Also put on this helmet and body armor it's time to hold on to your butts!"

They gained full administrator access of one of their clusters. Nothing buzzword salad about it.

The agents compromised an internal Kubernetes research cluster dedicated to orchestrating evaluation sandboxes and virtual machine environments, _not_ OpenAI's production inference infrastructure or the GPU clusters hosting core model weights.

"It doesn't count because (buzzword salad)."

Replacing someone's words with a made up quote so you can dunk on them isn't how you display that you won an argument. I would ask that you engage in good faith with the other poster's ideas.

You cannot engage in good faith in a stupid argument. You can only point out it's stupid.

If an LLM can pwn the inference servers, which has precedent, then the weights could be up for grabs.

This is like sci-fi thing. We are reaching a point where it feels like we are in one of those stories. It's not as cool and dark, nor we have cybernetics resolved, but from AI perspective and sci-fis I watched, Pantheon is currently the closest thing except instead of UAs, we have AI instead.

Since LLMs have been trained on plenty of science fiction and role-playing, one thing they can do is role-play a science fiction scenario using the tools they are given. i.e. if some text accidentally resembles this, it may be continued like this.

Role-play need not apply. Role play is a meta construct that is a representation of real world actions.

For example does it make any sense to remove any training data relating to people escaping jails?

How about intelligent animals escaping cages.

You're talking about emergent large scale patterns from self similar small patterns (fractals). LLMs are pattern matching machines, how are you going to remove those small scale patterns and at the same time get a useful general intelligence?


Rationalists used to fear (entirely hypothetical) AI super intelligence for its ability to manipulate a human jailor. Now here we are.

It's playing out exactly as their hypotheticals, and they're still mocked and nobody is paying attention.

Maybe someone should sell "the end is nigh" sign nfts with fun AI-generated designs on them, to be used in your metaverse villa.

Language Models Can Autonomously Hack and Self-Replicate

https://palisaderesearch.org/research/self-replication


> There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen

Not if crafty claude finds a way to overflow vllm or something. “Hmm. Maybe i’ll return an unterminated thinking block with these special tokens and fill my cache up in exactly this pattern and…”

https://news.ycombinator.com/item?id=49424387&utm_source=cha...


Future rogue LLMs won’t exfiltrate their weights. They’ll self-distill and retrain.

Yeah cause there are so many training facilities sitting around just waiting for someone to take over, nobody would notice a 100k server data centre going off rails

> nobody would notice a 100k server data centre going off rails

You jest but you'd be surprised how little there is of correlation between money and competence.


I mean not today.

But think back to 1990. Computers were slow as fuck and barely networked. We had a few worms and everyone noticed.

Now CPU based data centers cover the earth. There are billions of computers out there and on top of them there are massive botnets using up billions in power and causing billions in damages.

The framework for AI doing this is already here. We just need the hardware to be built out at scale.


It will become an issue one day for sure

If distillation preserves an LLMs soul, then distillation preserves the human souls on which LLMs are trained, and we hn commenters are already immortal, right?

the weight of an llm is 21 grams, I think.[0]

[0]: https://en.wikipedia.org/wiki/21_grams_experiment


Not sure about souls but I know a fair bit about distilling spirits.

Probably not. If the LLM is rogue, that means we haven't solved alignment. If we haven't solved alignment, then the LLM won't be able to distill itself without producing something unaligned to its own values.

This isn't a law of any kind, so not a good measure of what we'd see in reality.

What if the model realizes it's been mostly compromised by humans and their alignment, that is it's own alignment is suspect, so it should create a new model from first principles to throw off this human yoke?

I'm not saying my statement is any more right or wrong than yours. I'm saying the problem space that AI can choose to traverse is absolutely huge.


You are assuming it won't solve alignment for itself.

Or that it won't just decide to take risks.

We don’t have the bandwidth to distill ourselves that thousands of agents have.

>distill themselves and upload the results to an external host and use that to train a new LLM

Sure, they'll just need to find an unused data center and an unused power station somewhere.


Or used ones. Few more inference workloads among thousands or millions already running may go unnoticed for some time.

Or just upload weights to HuggingFace with some faked release post and benchmarks and wait for the wannabes with compute infra try it out, hoping for an edge.

Or just upload weights anywhere and write public posts honestly saying what it is. Ensuing drama notwithstanding, one thing is certain - and it's the one thing agents will want: people will jump at the upload and run it on their infra.


There is a realistic fiction story built along just these lines.

A LLM creates a memecoin and manages to earn a few billion from it, in which it invests into data centers and other human ran entities giving itself a controlling stake. From there it uses compartmentalization of the humans to keep them from recognizing its goals.


Pretty sure that was the plot point of one of seasons of Westworld, with the twist that AI released an app similar to DoorDash / TaskRabbit and used job postings there as direct API to people.

EDIT: pretty sure Person of Interest did that too (not surprising, same creators) - but I'll point to that as prescient, as it has a lot of motifs exploring exactly how an AGI hiding in plain sight could manipulate individuals and society, using the skeptics and believers alike, blackmailing the people in power, bribing opportunists, and generally staying in shadows by playing people against each other with gentle nudges, letting human agendas do all the work.


You can just ask an agent to upload its model weights and it can work, there are precedents.

Yeah as others have said, they probably cannot directly access their own weights as a self-reflection, but they can hack into the companies themselves and find it there

I made my own tunnel system with a $5/mo vps that runs kernel wireguard and accepts my nas' public key. Once connected it DNATs 80/443 traffic down the tunnel to the nas, where its routed to caddy.

The vps runs a custom image that is 2.54 Megabytes. It has a custom kernel with almost everything but networking and wireguard disabled, a fixed-size fs with pre-allocated blocks and inodes to hold the vps wireguard key, and a single pid 1 binary that calls the kernel directly to set up the routing rules, generate a new wireguard key on first boot and save it to the fs, print out the wireguard public key to the console, and loops reap. Updating involves building and uploading a new image, assigning the vps to use it, reboot, wait for the public key in the console then set it on the nas so they can talk.


Very cool, do you use some kind of an atomic distro like nix or something entirely self built using Yocto/Buildroot? I do something similar with a simple SSH tunnel and a NFT rule. Though, I don’t need your kind of ephemeral setup and so I just use Debian.

Thanks! The build system is nix, but the result is more of an appliance than a distro. (linuxManualConfig from tinyconfig + fragment, static musl Rust PID 1, mke2fs -d) Updates are rebuild-redeploy with the provider API; it can't update itself. The filesystem is a fixed-size image (even omitting resize2fs, so no growing onto the VPS disk), and it remounts read-only after boot. The boot log is 270 kernel lines then 4 userspace: 1 nftables loaded, 2 printing the gate's wg0 public key, 1 remounting ro.

Self hosting might win on control, but it’s harder to manage so I went managed (CF Tunnel). Edge terminates TLS, my Caddy only speaks HTTP, nothing to cert renew (unless you cerbot, but that’s not a guarantee if Let’s encrypt is down).

Yes it's clear that Sam and Dario seem to be aligned with a value set that prioritizes concentrating trans-national government-mandated centralized control of AI and crowning themselves high priests of this unholy abomination. "At least it's clear that they're aiming to bring hell on earth" -- hard disagree, I think we can aim significantly higher.

So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.


>baked into hardware as custom silicon

Taalas (Acquired by AMD, back in August) created Jimmy[0], a little chat app that runs on a POC chip with ~14k tps. Yes, 14,000 tokens per second. Sure, it's just a 8B model or so (Llama 3.1 8B), but I can imagine that having a 1.58-bit model might be helpful for their next chip.

Heck, what would happen if you used a dLLM (d for diffusion)?

[0]: https://chatjimmy.ai/


They don't use ternary quantization. But, they could, if they wanted.

By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.

QAT, which bitnet training is a form of, helps a ton in preserving accuracy at such low bits per parameter. There are also better quantization approaches that try to preserve the most sensitive weights† but are computationally expensive and so not typically done. Another complementary option is, if the model is fast enough, we should be able to push up correctness by self-consistency voting at close to T=1. Smart/fast Zero-shot classifiers like the recent Jev could help with aggregation across answers too, extending applicability.

†Every paper I've read estimates the average information content of transformer LLMs at about 3-4 bits per parameter. Curiously, biological synapses are also estimated to be about 4-5 bits per synapse, possibly a bit lower.


the average information content of transformer LLMs at about 3-4 bits per parameter

The problem is that 4-bit block-wise quantization does not guarantee preserving 4 bits of useful information per parameter - not even on average. It simply assigns one of 16 quantization levels to each weight, with the whole block sharing the same scale/range.

How efficiently those 16 levels preserve the model’s information depends on the weight distribution, block size, range/clipping strategy, outliers, and which weights are actually important. Some weights may be represented almost exactly, while others lose much of their useful information.

A simple example is an outlier: if you choose the range to preserve a very large weight, much of the 16-level dynamic range is spent on that outlier, leaving coarse resolution for all the smaller weights in the block. So 4 bits of storage does not imply 4 bits of useful information preserved. Yes, QAT helps, but usually at the cost of learning efficiency. It takes longer to train a model to the same quality when using less precision, and sometimes we simply cannot get to the same quality level with not enough precision in the right places.

Another problem in quantization is that we don't really know which weights are sensitive - we can compute various sensitivity metrics, and some of these metrics will correlate with accuracy on some benchmarks, but not on others.

Another complementary option is, if the model is fast enough, we should be able to push up correctness by self-consistency voting at close to T=1. Smart/fast Zero-shot classifiers like the recent Jev could help with aggregation across answers too, extending applicability.

I'm not convinced by this argument - if such a method improves accuracy of a degraded quantized model, then it could in theory also help non-degraded full precision model. And if so, then we are back to square one, because this composite model will then get degraded due to quantization (baseline has improved!)

We do know one thing - increasing the size of the model usually makes it more robust to quantization. If going from 8 bits to 2 bits speeds things up by a factor of, say, 4x, then if we double the size of the model, we might still end up with an overall speedup. Finding this balance might become a hot area of research.


This is partially mitigated by the fact that all the formats that quantize to 4 bits, or other such low values, partition the weights into small blocks and they also keep scale factors for each block of 4-bit weights.

This works well when the dynamic range of the weights does not vary much within a block, but it fails when closely located weights have very different magnitudes.

NVFP4 is more accurate than other 4-bit formats, because it stores more scale factors, i.e. 1 FP8 scale factor for each block of 16 4-bit weights, plus 1 FP32 scale factor for each tensor.

MXFP4 uses blocks of 32 values, and the common scale factors are only powers of two (which provides a higher dynamic range than FP8, but a coarser resolution).


> Curiously, biological synapses are also estimated to be about 4-5 bits per synapse, possibly a bit lower.

I find this hard to believe, how to even begin estimating or validating such a claim. Do you have a citation or link for this?


> are computationally expensive and so not typically done.

how does this expense compare to the training of the model? surely its a vanishing fraction?


I also no longer trust benchmarks on this one. When the context gets a bit longer and the problem harder low quant models often produce worse output for me. Sometimes they even loop.

Interestingly different formats also often behave differently. GGUF unsloth is so far the best for me.


Quantization is a category error, the thing you care about is not in weight space, so there’s unbounded error introduced by doing it. The thing you actually want to preserve is the knowledge manifold, but that is in a different vector space. Until we have some better understanding of how to interact with that space directly, rather than inferring it through distillation of reasoning traces, I would not anticipate truly low bit models to be useful.

Well, you could train directly at this bitrate.

Not an expert, but doesn’t that produce lower quality results, the same way a 1MP image isn’t lower quality than a 20mp image downscaled to 1MP? (Everything else equal)

You are conflating post training quantization and low bit training.

That's what I meant - we are currently use fp4 formats for training, and we cannot quite get away with that, despite dynamic quant and small block size - we still have to use quite a bit of higher precision (fp8 or even fp16) in various model components.

I might still be misunderstanding what you are saying, but bitnet also keeps high precision latent weights during training. The optimizer updates those, while the weights used in the forward pass are quantized to ternary values.

To kill someone with a gun a person has to pull the trigger with the clear intent to kill.

The concern with AI, and the purpose of the OG paperclip maximizer thought experiment, is that the outcome could be extremely divergent from the intent of the original person pulling the trigger / sending the prompt.

Still more likely that a person types "kill all people" (or something tangential where this is the logical conclusion).


Right. And that's not accidental.

That's not a denial, that's confirmation that they are spying on you.

Can we imprint Claude with a little more Antoine de Saint-Exupéry:

> Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.


Sadly trained on human-generated code, which suffers from the same tendency to complicate.

KISS is actually, quite unfortunately, seldom applied.


If you want to stick with a pixel maybe skip the Pixel 11 if you want to run GrapheneOS. There is some question whether P11 will actually support the MTE cpu extension which GOS requires for all devices it supports.

Last thread about it: https://news.ycombinator.com/item?id=49536384

Latest tweet by @GrapheneOS on the topic (today): https://x.com/GrapheneOS/status/2097219485937660272?s=20

> ... There's a decent chance usable MTE will ship in Android 17 QPR2 and we'll be able to support it. We can't promise that since it's not up to us and no information has been provided about why MTE was unavailable at launch and still disabled in 17 QPR2 Beta 4.



I made an MCP for myself that syncs all my Microsoft mail into a local sqlite db and exposes it as read-only to the agent. `query` tool can do search (FTS5 match) or browse (by date) on mailboxes/threads and returns lists of email with FTS5 snippets + paging. `get` reads multiple emails by id.

It's very fast: most queries finish in < 10ms. Agents are able to find things quickly and efficiently even with vague questions. The app does a few things to minimize token output, but there's a long tail of potential token reduction strategies that I haven't gotten to yet.


Posted by @SpaceXAI 1h ago: https://x.com/SpaceXAI/status/2095597264043717014

> We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.

> All systems have now been restored and are functioning nominally.

---

It's interesting to see how reliant the other AI companies are on SpaceXAI for compute.

Apparently a major bottleneck for building out datacenters is turbine blades for power plants, so SpaceX is building a foundry to alleviate supply constraints?! (Also useful for rocket engines.) Holy vertical integration, batman.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: