Hacker Newsnew | past | comments | ask | show | jobs | submit | potus_kushner's commentslogin

dude talks about measuring C while the code he shows is C++.

The code is in _C_ the syntax highlighter says _C++_. _C++_ highlighting works with _C_.

A number of RTOS allow the use of restricted C++ with complex features disabled; often called C with classes. The compiler used will most likely _C/C++_ versus only _C_. Example: µC/OS-II used in military, aerospace, and medical applications. [0]

[0] https://micrium.atlassian.net/wiki/spaces/osiidoc/pages/1638...


especially if they get away unpunished for their price fixing. https://web.archive.org/web/20260701064235/https://www.polyg...

at first this looked like something one could run on CPU with 64GB RAM with a 2-3 bit quant, at possibly half the speed of 3.6 35B-A3B, however the 50B ngram sidecar makes it impossible. and oddly, unsloth's page lists the ngrams as 50GB even though they say it's in 4 bits. should be 25GB according to my math. anyway, the new ngram architecture makes it pretty much unusable for regular folks who cant afford more than 32-64 GB ram in this RAMocalypse.


I suspect we will see optimizations where the various vectors of the n-gram you actually use are hot in vram, the rest are warm in system memory and then cold storage on nvme. Same with MoE. If your workflow is particularly same-y then you're looking at cache miss below 5% with NTP/MTP turned on and the right harness. Agentic "openclaw" type stuff cache miss might be below 1% in the right local llm setups. There's been zero exploitation of n-gram stuff yet, it will be very interesting as things progress.


the xhigh version looks amazingly good for a 2 bit quant.


that's why they're called "the company" in insider circles. they still do, openly and probably also covertly. their investment arm is called in-q-tel https://www.iqt.org/portfolio , funding things like gitlab and docker. they probably run half of silicon valley. the dead giveaway is a funding story about 2 guys in a garage or dorm.


The garage or dorm story is the biggest myth they keep repeating, it's basically nonsense.


another strong Qwen 3.6 35B-A3B fork with open weights.


impressive work. but a GGUF release + (ik_)llama.cpp PR would be appreciated since not everybody has a mac.


Thanks. I agree, GGUF and upstream contributions are on my radar.


Can't you convert?


pretty impressive model for its size, if the benchmarks can be trusted.


hopefully for us mere mortals without $4k+ hardware a 35B MOE model will be released. or a new prism ternary bonsai model based on this one.


the ramblings of a madman.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: