Hacker Newsnew | past | comments | ask | show | jobs | submit | Eliezer's commentslogin

> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks

If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.


You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. It's a massive problem with human authors too, which is why they do a lot of lorekeeping and editing, so the AI should be afforded the same tools if we are debating human level ability.

I agree on the one-shot (which is not a fair comparison because nobody oneshots a good story), but I'm not convinced this part hasn't reached AGI already.


> You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes.

Yes, and I could script a truly marvelous proof if this textarea were but a little larger :)

Hand waving doesn't count for much these days when you could spin these things up quite quickly to prove the point, so the GP's claim seems much stronger than whatever you're not convinced of?


This is the first time I've been accused of not spinning up an app at my expense to win a meaningless argument on the internet.


Why on Earth would anyone take anything this man said to be informative about anything, other than what he thinks will currently sound best that minute to the current Zeitgeist that day?


Just because he's a professional bullshitter doesn't mean he shouldn't be called out every time he waffles.


Safely aligning a powerful CEO is difficult.


He's already proven repeatedly that he'll break out of containment


You can turn that off in settings.


The safety teams are trivial expenses for them. They fire the safety team because explicit failure makes them look bad, or because the safety team doesn't go along with a party line and gets labeled disloyal.


The first customer is always the investor these days, so anything that threatens the investor's confidence is bad for business.


Came here to say the same.


Every time somebody writes an article like this without any dates and without saying which model they used, my guess is that they've simply failed to internalize the idea that "AI" is a moving target; nor understood that they saw a capability level from a fleeting moment of time, rather than an Eternal Verity about the Forever Limits of AI.


Funnily enough we have had those comments with every single model release saying "Oh yeah I agree Claude 3 was not good but now with Claude 3.5 I can vibe-code anything".

Rinse and repeat with every model since.

There also ARE intrinsic limits to LLMs, I'm not sure why you deny them?


There's intrinsic limits to vanilla transformer stacks. Nobody knows where they are. We don't know how unvanilla Opus 4.6 or GPT 5.3 are. We don't know what's in development or which new ideas will pan out. But it will still probably be called an "LLM".


Exactly. Basically their thought are invalidated often quicker than they hit the publish button.


"2024 AI was a horse". People really like to imagine that the last 6 months constitute their true observation of the new eternal state of the future.


Exactly. We're headed for a discontinuity, not an inflection point.


It would not surprise me at all for bb7 to exceed Graham's number. Just a Kirby-Paris hydra or a Goodstein sequence gets you to epsilon zero in the fast-growing hierarchy, where Graham is around omega+2.


The 79-bit lambda term λ1(λλ2(λλ3(λ312))(1(λ1)))(λλ1)(λλ211)1 in de-Bruijn notation exhibits f_ε0 growth without all the complexities of computing Kirby-Paris hydra or Goodstein sequences. Even that is over 60% larger than the 49-bit Graham exceeder (λ11)(λ1(1(λλ12(λλ2(21))))). I think one should be quite surprised if you can climb from f_4 (2↑↑2↑↑2↑↑9) to f_{ω+1} (Graham) with just 1 additional state.


Thanks to OpenAI for voluntarily sharing these important and valuable statistics. I think these ought to be mandatory government statistics, but until they are or it becomes an industry standard, I will not criticize the first company to helpfully share them, on the basis of what they shared. Incentives.


Privacy.com has always been a working general solution to this, for me. Disposable CC aliases with spending caps.


Privacy.com won't provide lawyers to you when Oracle sends letters of demand to you for unpaid invoices.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: