Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
from
login
P(Kill-Switch|Detection)
(
lesswrong.com
)
1 point
by
kp1197
3 hours ago
|
past
|
discuss
Watch AI materials-science and bioscience abilities closely
(
lesswrong.com
)
39 points
by
joozio
11 hours ago
|
past
|
49 comments
Another Slice of Swiss Cheese for Untrusted Monitoring
(
lesswrong.com
)
2 points
by
joozio
15 hours ago
|
past
|
discuss
Astra and Fable still hack on simple variants of alignment evals from 2025
(
lesswrong.com
)
464 points
by
Levitating
1 day ago
|
past
|
227 comments
The Talker Does Not Control the Doer (In Current AIs)
(
lesswrong.com
)
1 point
by
jstanley
1 day ago
|
past
|
1 comment
An interesting anecdote from our Hacker Opus work
(
lesswrong.com
)
2 points
by
yurivish
1 day ago
|
past
|
discuss
How My Students Think About AI
(
lesswrong.com
)
51 points
by
paulpauper
3 days ago
|
past
|
10 comments
Adaptive Agentic Worms Are Here
(
lesswrong.com
)
2 points
by
speckx
5 days ago
|
past
|
discuss
Astra and Fable still hack on simple variants of alignment evals from 2025
(
lesswrong.com
)
2 points
by
yurivish
6 days ago
|
past
|
discuss
Interpreting GPT: The Logit Lens
(
lesswrong.com
)
2 points
by
Bluestein
6 days ago
|
past
|
discuss
From safety research prompt to cross-model universal jailbreak
(
lesswrong.com
)
2 points
by
gmays
7 days ago
|
past
|
discuss
My Students Think About AI
(
lesswrong.com
)
5 points
by
alphabetatango
10 days ago
|
past
|
1 comment
Asking agents to make money to survive
(
lesswrong.com
)
4 points
by
paraschopra
10 days ago
|
past
|
discuss
What is nueralese and why is it bad
(
lesswrong.com
)
77 points
by
tristanMatthias
10 days ago
|
past
|
59 comments
How concerned should we be about Astra's recurrent architecture?
(
lesswrong.com
)
151 points
by
yurivish
11 days ago
|
past
|
129 comments
METR Researcher Thomas Kwa Hired by OpenAI
(
lesswrong.com
)
1 point
by
qlte
11 days ago
|
past
|
discuss
The Library of Scott Alexandria
(
lesswrong.com
)
4 points
by
benatkin
11 days ago
|
past
|
1 comment
Models may behave differently in graded episode
(
lesswrong.com
)
2 points
by
ddp26
12 days ago
|
past
|
discuss
AGI and the Efficient Market Hypothesis (2023)
(
lesswrong.com
)
2 points
by
Metacelsus
13 days ago
|
past
|
1 comment
How My Students Think About AI
(
lesswrong.com
)
5 points
by
pella
13 days ago
|
past
|
1 comment
P(Kill-Switch|Detection)
(
lesswrong.com
)
2 points
by
kp1197
14 days ago
|
past
Starting AI Safety Study Group to Do Arena Curriculum
(
lesswrong.com
)
2 points
by
joozio
14 days ago
|
past
Cooperating with aliens and AGIs: An ECL explainer
(
lesswrong.com
)
3 points
by
Bluestein
19 days ago
|
past
Prompt Sufficiency: A Missive for the Managerial Class
(
lesswrong.com
)
1 point
by
kp1197
22 days ago
|
past
We Must Remember That Our World Contains Hell
(
lesswrong.com
)
1 point
by
paulpauper
24 days ago
|
past
LLMs are (still) mostly powered by imitative learning, not RL
(
lesswrong.com
)
3 points
by
wslh
24 days ago
|
past
Can an LLM make a feature-length movie on its own?
(
lesswrong.com
)
2 points
by
mchinen
26 days ago
|
past
Recursive Middle Manager Hell
(
lesswrong.com
)
5 points
by
rzk
27 days ago
|
past
LLMs are (still) mostly powered by imitative learning, not RL
(
lesswrong.com
)
1 point
by
surprisetalk
28 days ago
|
past
Kimi likes causal decision theory more after RL in twin prisoner's dilemmas
(
lesswrong.com
)
1 point
by
0xkato
29 days ago
|
past
More
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: