Hacker Newsnew | past | comments | ask | show | jobs | submit | c0rruptbytes's commentslogin

> This argument doesn't work for countries which don't care what their citizens want, like China.

you didn’t have to use China as an example, the US clearly does not care what its citizens want as the most popular policies are never even discussed or proposed in congress

meanwhile, China destroying their housing market to decommidify it so everyone can have housing…they seem to care about their people more


another obvious example is climate change, where China is way more committed to solving it than the US. the division of the world into good guys and bad guys by a rather arbitrary definition of democracy looks so naive, even stupid. Anthropic is not a Democratic institution either, should we trust it? if this proposal for pacing the frontier is to be taken seriously, Dario should give its competitors, also international ones, the same assumption of good faith that he demands from us.

the iphone mini unfortunately did not kill nor do most small iphones

are they releasing the weights too?

open models don’t imply local models… it just allows more people to run them, most likely with nvidia GPUs

so many inference project, omlx already supports all of this and has a 1000 people trying to optimize it constantly

Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX

Interesting! I'll check it out

m5 max really fixed pp with the better matmul support, im sure the m5 ultra will be even crazier

the sparks have much slower memory bandwidth is the trade off


I believe the dgx spark is still twice as fast at prefill as the m5 max, but the ultra should get closer to parity.

Another benefit of the 2x spark setup is that you can parallelize to ~6 streams pretty efficiently.

All depends on the workflows you’re using it for.

I’m quite excited for the M7 class machines.


the rr suite seems much better for that


Of course you can use the native UI of all the apps in your ecosystem, the biggest feature of Hermes for me personally is that I can run any task in any of my 30 or so self hosted tools from a single chat interface (matrix), which is also quite secure. No longer do I need 30 open tabs and lots of clicking around, one sentence in my favorite chat app (even on the go in the phone), and many tasks can be executed at once. Unification of control.


The same concept works with the arrs, too, doesn't it?


OpenAI measures their internal token usage in “rolexes” - it’s literally a flex to be a token burner

i can imagine insane amount of capital is wasted on these two companies compared to the efficiency elsewhere


And despite the enormous capital expenditure, Chinese models are nipping at their heels at what must be a fraction of the cost. Sometimes constraints are healthy for inducing creative solutions.


it's for me


The 512GB could run GLM 5.3 which is Opus level


GLM 5.2 in NVFP4 is 465 GB. It would be a tough fit.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: