Hacker Newsnew | past | comments | ask | show | jobs | submit | JCharante's commentslogin

Maybe they (oai) want to pump their marketshare on openrouter lol


Any benefit of pumping marketshare on open router?


if I was going to IPO I'd want higher marketshare and openrouter is the most popular way to measure model usage (not saying it's the most accurate b/c it's not)


most mice require more steps per mile of distance


I wouldn't say it's the vast majority if Tokyo/Chiba/Saitama/Yokohama holds like a third of the population. Add in those other cities you mentioned and the majority of the population lives in cities.


Of course in any developed country the majority of the population lives in cities. What I mean is that Japanese cities, outside of a few core areas, are really ugly and depressing.


just don't live in the countryside then, there's a reason their population is decreasing


People are forced to go to the big cities for economic reasons and not lifestyle reasons. Unlike America, it’s not just white collar office workers who must relocate to urban centers. This concentration forms a feedback loop.


Unfortunately my workplace is in the countryside. Someone decided it was a good idea to build a big international research center in the middle of nowhere. Which is now a liability for the project, as few people want to come work here.

https://maps.app.goo.gl/6Qf8EGk97jStkJFP7


I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead.


Anecdotally, it feels like Opus, Fable, and Sol "get distracted" when you use them for writing code. Great at reasoning and coordination but they will go off on a tangent and refactor half the code base. I only use them for reasoning (of course) and coordinating subagents.


mind sharing hints/links on your harness/flow setup?

I did several attempts with naive prompting, but spent more time babysitting than actual flow


Just an anecdote but thats why Deepseek v4 flash 0731 is my current favorite model. It's really not very "eager" and stays on the task at hand.


Have been feeling the same. There's a sweet spot that threads the needle between "too dumb to search the right thing / relay the correct results" and "too smart to just stop overthinking and just report the damn thing"


founder of castform here again - slightly unrelated to retrieval but on the topic that folks are discussing here, i was actually collecting benchmarking various coding traces for the purpose of training a model router and surprisingly, luna held up very well against sol and terra. it was able to solve close to >95% the that sol can handle at a fraction of the cost. have not benchmarked the OSS models yet but will add the popular ones to the list like Deepseek Flash and Kimi k3 to see how they fare.

will share the full results soon!


we actually have the test benchmark against luna too! it's just not in our title but you can see it in the first diagram below the title. luna does pretty well tbh but sol is just a tad bit better. but luna is way cheaper.

if you want to dive down into the various traces of the benchmark, you can check this out: https://app.castform.com/train/a7a898f6-d802-4908-b044-acb81...

- founder of castform


Any data or public links you can share? That surprises me


surprised running paxel doesn't disqualify you, the program sounds sus


I wonder what the tipping point is where machines require less energy than humans to accomplish the same tasks.


I used to donate blood regularly but now that I'm in Japan they require me to be decently fluent in Japanese to "understand" the risks, despite having done it a bunch of times in other countries (and other medical procedures not requiring Japanese knowledge).


I mean eve frontier is a pretty huge fork from almost scratch, did they backport changes into normal eve?


I can't say I have a definitive answer, but there is a blog post from 2025 that seems to suggest they haven't yet: https://nosygamer.blogspot.com/2025/05/fanfest-2025-upgradin...


you don't have to arrest everyone, just make the likelyhood of being arrested non-zero to deter people


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: