Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
kqr
on Aug 12, 2025
|
parent
|
context
|
favorite
| on:
Evaluating LLMs playing text adventures
Ah, I see what you mean. Yeah, there was too much output from too many models at once (combined with not enough spare time) to really perform useful qualitative analysis on all the models' performance.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: