Hacker Newsnew | past | comments | ask | show | jobs | submit | zbrock's commentslogin

Correct on the first part, partially correct on the second. LOC is a bad metric, but it is at least a legible one. Lots of people working on better ways to measure Software Productivity!


At the time we wrote the article we hadn’t released the product and weren’t ready to talk about it. It was an internal prototype that looked very much like the current Codex app.


So, did this internal prototype ultimately end up being used to create/influence a real product, e.g. Codex app?


Yep!


No detail about this in the article or your comment here, but the voluminous lines of code get a big call-out. Very interesting!


Hello! I’m one of the three engineers who write this piece. Happy to answer questions.


Interesting write up!

Have you been able to extract libraries or tools from this project yet? If so how was that experience?

That is, do you see yourself releasing a metric harness, or sub-projects that are equivalent of ActiveRecord, zod, or similar open source tooling that frequently originate in a large in-house project - and then is exported out as a stand-alone toll, utility, library or framework?

Because while ai can reimplement minor tools, it's utility entirely depends on the existence of solid tools, libraries and frameworks.


Very cool article!

- are other teams adopting this approach? What’s the blockers if not?

- have there been problems where the models alone were not enough to debug and the devs had to fix it themselves?

- as the rate of changes has increased with more devs how have you dealt with concurrent writers with merge conflicts?

- if there was anything you could change in the approach you started with, what would it be?


1. Yes! Many teams internally have adopted a lot of the same practices we outlined in the blog post. Ryan has also been spending time both internally and externally helping companies figure out how to do this in their code bases.

2. Hmm, kind of. There have definitely been issues the models can’t one shot. But we still use Codex to write all the actual code with human guidance.

3. More agents :) Some teams are experimenting with centralized Agent mediated integration queues, others use normal merge queues, many have local Codex threads that monitor CI to resolve and land conflicts or failures.

4. Today’s models and codex app. We started doing all this with gpt-5 and codex-cli. The tools today, 9 months later, are so much better than what we had then.


Have you built any tooling or products around all of this and deploying it somehow? I’d love to learn more and share notes, because we’ve been doing this too. About 3100+ PRs merged across our 4 person team in 4 months. Impossible without harness engineering, and I agree, the tools are getting even better.


If you were paying commercial token rates, what would the cost have looked like?


Fantastic job!

Can you share what type of project that was? On the spectrum from a database engine to cat picture sharing web site (very high demand for correctness vs very lax).


This was primarily an Electron app with some small hosted backend services


Have you been satisfied with the quality of code generated by the model? Or did you have to tweak some rule file or skill to improve it? Or is human-readable code not even a goal at this point?


We spent a lot of time tweaking skills, doc files, and prompts. I’d say that was our primary activity as engineers. Our job became tweaking the harness every time we got code or results we didn’t like. Eventually we were pretty happy with most agent runs, but we were always happy to just throw out ones that didn’t meet our standards. I think more than half didn’t.


Were those em dashes you, or GPT


They did that already in the 80s, they're called Incentive Stock Options


"What are the current salary bands at your company for someone with my level of experience?"


We sync your items, settings and payment history between multiple devices. That was...not easy.


Nothing worthwhile is ever easy. =)


Thank you! Our in house video team is pretty awesome.


Do you know of any high traffic sites that aren't?


We pair a fair bit at Square. We require that all code either be paired on or go through code review, but let each team decide their own process. Some of the teams pair 90+% of the time, and some pair only infrequently.


San Francisco, CA

Square (https://squareup.com/jobs)

We're going to change the way people pay. We've got some big backers (http://www.crunchbase.com/company/square), a bunch of really smart people and a product that people absolutely love.

Email me (zach -at- squareup.com), or apply through the site if you're interested.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: