Hacker Newsnew | past | comments | ask | show | jobs | submit | hgoel's commentslogin

Yeah, the latest batch of <256GB Chinese models are really nice. They're far less cryptic than Claude, and competent enough to feel almost near Opus. I canceled all my subscriptions and switched to running the Chinese models locally (not as a cost saving measure).

As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago (edit: despite being small enough to usably run on prosumer hardware). It seems clear parameter efficiency can still be improved dramatically.

In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).


Especially after the blatant cover up of their uncontrolled bot swarm infesting the internet, and the feckless "hopefully we do better" response upon being caught, I don't think OpenAI deserves much grace until they properly explain themselves.

We had all assumed that surely the supposed smartest engineers in the world, with access to the most computing and a direct view of model capabilities, would take sandboxing and cybersecurity much more seriously than they have turned out to do. It follows that while we might assume they take user data privacy seriously and have tight controls on who can access it, it's possible they do not actually do that.

At this point any initial trust is dead and has to be re-earned.


They can't exactly occupy the same space due to the Pauli exclusion principle. IIRC that's believed to be the final "barrier" that prevents neutron stars from collapsing into black holes. But this is also getting into the tricky parts of wave particle duality, so the precise details are a bit difficult for me too.

Only cats can violate Pauli's principle. Anybody who had a cat can attest to the fact they can go through walls. Just close a room with a cat inside and - given enough time - the cat will escape.

So Schrödinger’s cat has an orthogonal state!

(not inside the box)


Actually, the third possible state is Bloody Furious.

How could you measure that the cat was inside the room?

Thermal residue?

GCN was such a promising compute architecture, AMD even pioneered stuff like async compute and compute shader heavy rendering pipelines, only to never seriously go beyond that on consumer gear.

I agree with your assessment that the story of supporting the competition, only to get burned, has repeated many times with AMD outside the public eye. It's why I don't put much stock in claims that things work great as long as specific flags are used.


Social media is already full of hentai art of that sort and that isn't a problem to most countries, so a little thought leads to the obvious point that your basic assumptions are incorrect.

Minors were involved and people were generating sexually charged images ("donut glaze" etc). This is why other countries, like South Korea, also went after X under deepfake related regulations.


Prompting for some of the results was almost the "Computer, do a breakthrough. Make no mistakes." meme. Just someone telling the model to keep trying a couple of times.

Unfortunately we don't actually know what kind of prompting was done for the more prominent results.


It’s definitely not how I work, I’d need to read the model responses and set a direction for the model to go

Same here, it's one thing if I'm just screwing around, but if I'm trying to do anything serious, I need to at least have a handle on what it's doing, and, when thinking traces are available, keeping track of any logical errors in the model's reasoning.

you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like:

7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.

Citation: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3..., section 6.2.2

You're not going to get a handle on what it's doing. The thinking traces are there to make you feel better about yourself.


Anthropic specifically calls out that example as

>"illegible reasoning in a few reinforcement-learning environments over long rollout"

Yet, I get the point that you're making: those tokens essentially are an internal scratchpad for the LLM which isn't required to logically lead to the output.

This video presentation of the paper you linked was interesting: https://www.youtube.com/watch?v=hUp3zh23aHw


I'm referring to models where the actual traces are available. Eg. I've been using a local Qwen3.8-Next-Flash lately.

Still not meaningful -- https://arxiv.org/pdf/2504.09762; even for local models, the reasoning traces are often filtered and summarized to sound sensible to humans. And even if not, they don't necessarily represent what the model is thinking.

Hmm interesting, thanks for the link, I'll have to give that a read.

Why is the model playing poker

It depends on the time the job itself takes. If you're having the LLM handle a training run for another model, the LLM is probably spending most of its time waiting for iterations rather than consuming tokens.

For a task I left a local model running on overnight, only ~100k tokens were used because most of the time was just waiting on tests to finish, then waking up, tweaking a few settings and trying again.


Hi!! This is basically why I built Observer (https://github.com/Roy3838/Observer, FOSS). It runs a small local model that watches your screen/terminal and pings or calls you when the thing you're waiting for actually happens: tests finished, run crashed, agent stuck on a prompt.

So an overnight loop like yours wakes you only when it needs a decision, instead of you waking on a timer to check.

But in this case it would be triple LLM inception, one training another and a third one monitoring everything is done correctly :p


Apparently between technology and soul crushing authoritarianism at an unprecedented scale, you would much rather the latter than the former.

Same with NYC, "going to The City" usually just means going to Manhattan.

The city of London literally has a city also named London inside it and yeah, "Where are you going?" "The City" means clearly that tiny part. There is another small older city inside London, Westminster, but if you meant Westminster you'd say "Westminster" never "The City" in general.

I grew up equidistant from Philadelphia and New York. "The city" was exclusively used for New York regardless of the many cities between us

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: