if I was going to IPO I'd want higher marketshare and openrouter is the most popular way to measure model usage (not saying it's the most accurate b/c it's not)
I wouldn't say it's the vast majority if Tokyo/Chiba/Saitama/Yokohama holds like a third of the population. Add in those other cities you mentioned and the majority of the population lives in cities.
Of course in any developed country the majority of the population lives in cities. What I mean is that Japanese cities, outside of a few core areas, are really ugly and depressing.
People are forced to go to the big cities for economic reasons and not lifestyle reasons. Unlike America, it’s not just white collar office workers who must relocate to urban centers. This concentration forms a feedback loop.
Unfortunately my workplace is in the countryside. Someone decided it was a good idea to build a big international research center in the middle of nowhere. Which is now a liability for the project, as few people want to come work here.
I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead.
Anecdotally, it feels like Opus, Fable, and Sol "get distracted" when you use them for writing code. Great at reasoning and coordination but they will go off on a tangent and refactor half the code base. I only use them for reasoning (of course) and coordinating subagents.
Have been feeling the same. There's a sweet spot that threads the needle between "too dumb to search the right thing / relay the correct results" and "too smart to just stop overthinking and just report the damn thing"
founder of castform here again - slightly unrelated to retrieval but on the topic that folks are discussing here, i was actually collecting benchmarking various coding traces for the purpose of training a model router and surprisingly, luna held up very well against sol and terra. it was able to solve close to >95% the that sol can handle at a fraction of the cost. have not benchmarked the OSS models yet but will add the popular ones to the list like Deepseek Flash and Kimi k3 to see how they fare.
we actually have the test benchmark against luna too! it's just not in our title but you can see it in the first diagram below the title. luna does pretty well tbh but sol is just a tad bit better. but luna is way cheaper.
I used to donate blood regularly but now that I'm in Japan they require me to be decently fluent in Japanese to "understand" the risks, despite having done it a bunch of times in other countries (and other medical procedures not requiring Japanese knowledge).