Assuming the LLM never got anything wrong or otherwise had to be re-prompted, that means your devs were reviewing 130 SLOC per hour, on what was described as moderately greenfield (examining new implementations rather than comparing to old historical accidents).
How?
I don't want to sound flippant, but if the point is to add human thought to the mix, that's a high review rate even when examining small tweaks to an existing, working product, even with substantial AI help to pre-filter major gotchas before you bother spending a lot of human effort on the review. That's only 20-30wpm, but a review isn't just scanning or reading code, especially if you're trying to figure out how a new system which doesn't run yet will fit together.
Is that actually a high review rate? Especially if you know the language and domain. Sure, initially there's a learning curve for a new codebase structure, but lots of lines will also be trivial and many changes might also be similar to each other.
For small separate changes in isolation then maybe it's ok? But not for whole days 8 hours each.
But then you need to watch for bugs coming from interaction with previous changes and in 700k loc that might be nontrivial. How do you know which states are reachable and which are not? That takes time.
It only takes a botched condition here (forgot a "not"? swapped "and"/"or"?), a swapped variable name there, code that looks ok, but isn't.
Those numbers are in units of "Consumption per Real Dollar of GDP", which is defined as follows: "Calculated as energy consumption
divided by U.S. gross domestic product in chained (2017) dollars".
I don't know exactly what it meant by "chained", although from the context it does sound as though it might mean something like "inflation adjusted".
I am getting more and more peeved by using GDP as a single number. Cost of existing, aka housing and healthy food, skyrocketed compared to median income. But hey, the tech got cheaper.
It does mean they tried to eliminate inflation as a factor. In my experience though the basket of goods used for inflation calculations do a poor job representing the majority of consumers' and businesses' costs.
I don't understand the desire to remove Levent. "Off the clock, when Anthropic engineers want to break new ground, they use ChatGPT" sounds like a great ad.
All these small details will be forgotten in less than a year. But association "Navier-Stokes -- OpenAI" will be part of the history. In my opinion, they are thinking about long term PR here.
Buying 3 different sizes of a shirt because manufacturers refuse to standardize on sizing information and returning 2/3 of them seems like a standard practice for buying clothes online. Not sure why you think the stores don't account for the returns.
reply