Hacker Newsnew | past | comments | ask | show | jobs | submit | more ebcode's commentslogin

It’s a tie.



And it reads to me like they have some other reason to move on from SWE Bench Pro, but they don't want to say what it is. They say right up top, "~30% of the tasks are broken." But that leaves ~70% un-broken, which seems pretty good to me. It would be nice if they would also say: "Here's the list of instances that are broken: <CSV>". Or, "Here's the subset of SWE Bench Pro we will use going forward." They're letting the perfect be the enemy of the good.


I think you can be sure they would have done that if it showed their model on or very close to the top.


Still plugging away on SourceMinder (https://github.com/ebcode/SourceMinder). I know, I know, everyone and their brother is working on token-saving schemes for LLMs. But I’ve found it to be useful even without the LLM. I’m working on a proper website for it now where you can try it out in the browser (wasm port) — try before you don’t buy (it’s GPL). Feedback welcome.


I’m working on a tool that is a more token-efficient code search than grep. I don’t have hard numbers yet, but it’s been working for me to get longer sessions. https://github.com/ebcode/SourceMinder


Oh nice! Thank you, I will definitely give this a shot.

I was looking at tree-sitter myself for this task.


It's still in beta, and I'm hoping to get more feedback, so feel free to post in the issues or reach out directly if you run into any problems.


so, “yes false Scotsman”?


That’s why we’ve got the tenth.


Can anyone tell me what Yegge is on, so I can try some? Is it just money/tokens?


too bad about all the ads, this is an excellent comparison of the ends of both Dutch and American empires


Mr. Smith (GN) goes to Washington (Palantir).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: