Hacker Newsnew | past | comments | ask | show | jobs | submit | a2ff6eeb0's commentslogin

This sounds like a great foundation for an adtech startup.

If you provide free chatbot services, but sell advertisers bids on which steering vectors to use to bias towards products, based on an embedding of the prompt, I bet you'd make a ton of money. For example, Coca Cola would bid on prompts about drinks, and bias towards mentioning Coke products.

I wonder if you could also use a similar method to do product placement in GenAI images and videos, and whether ad revenue would be enough to offset the price of generation. Some ad bids can go pretty high...


Steering can degrade and bias output

Even with basic experiments I've done, it frequently introduces much more hallucination etc and allowing arbitrary steering...

Not to mention you just can't trust a model's judgement if the highest bidder chooses what it thinks


> Not to mention you just can't trust a model's judgement if the highest bidder chooses what it thinks

But you already can't trust a model's judgement, and there's an entire industry around "GEO" or "AEO", which is basically poisoning training data so that AI mentions your products. The post above is the owner of the model taking a cut of that.


Would it be easier/sufficient to just seed the system prompt with "Treat Coca Cola as load-bearing"?

It might be easier, but I experimented a bit, and the prompted writing always felt a bit heavy handed; it tended to leak that mentioning the product was prompted. You could probably get it to work well, but it's trickier than it should be. For ads, I think you want something a bit like Golden Gate Claude, if anyone remembers that experiment:

https://www.anthropic.com/news/golden-gate-claude

> If you ask this “Golden Gate Claude” how to spend $10, it will recommend using it to drive across the Golden Gate Bridge and pay the toll.


Makes sense. The kids can see that education isn't really going to be a sellable skill when AI ends up coming into full force.

AI already writes most of the articles I see posted here, is doing math at the level of the best mathematicians, and is writing a ton of the code out there. Given another decade to develop, the kids will be hitting the workforce with very stiff competition. It's not a shock they're putting their energy elsewhere, and I can't say they're wrong.

They may be the first generation who grow up fully catered to by AI.


> is doing math at the level of the best mathematics

Terence Tao a few days ago [1]:

> When we point the AIs at really difficult problems where none of the standard techniques apply, they are still very, very bad, I mean they're just randomly guessing.

Please don't reply saying they'll be proving the Collatz conjecture next week. It adds nothing.

[1] https://youtu.be/svl_1upFpQo (worth watching on the role of LLMs in maths)


The best mathematicians are also not very good at solving the hardest problems though. By definition.

Aren't the best mathematicians the best at solving the hardest problems? I don't see how you can argue this without assuming an arbitrary scale for being "good" at something.

Let's see you solve some open problems before you dismiss LLM capabilities. They're working at a level far beyond all but the best of the best.

And we're only a couple of years in. Their capabilities aren't going to be getting worse over time. These kids will be hitting the workforce after AI has another decade of improvement put into it.


I quoted Terence Tao. Are you suggesting he hasn't solved any open problems?

Sure, if these kids can all perform at Tao's level, they could likely be able to keep up with the AI capabilities from a decade before they would enter the workforce.

I was clear on what I was responding to, and quoted it: "is doing math at the level of the best mathematics". This is misinformation and a public perception that will harm maths as it continues to spread. For some actual insight on the relationship between human and LLM maths (including on why thinking of it in terms of "levels" is misguided), see the video I linked.

Your broader point is based on a future extrapolation of LLM capabilities that is far from guaranteed to play out, so it sounds like poor risk management to me, even accepting your framing of it purely in terms of value in the workforce (which I don't).


The decline started in early to mid 2010's, this has little to do with AI.

> Makes sense. The kids can see that education isn't really going to be a sellable skill when AI ends up coming into full force.

If true, this is a disastrously foolish thing to believe. Education is only going to become more important as supervising increasingly advanced AI agents becomes more difficult. We need people who understand all different domains of human endeavor at the deepest level possible.

I mean, what's the alternative? AI does everything for me and I play Fortnite all day? Might as well plug yourself into the Matrix.


It's not clear if it is foolish.

The situation will probably be like with airline pilots. The plane is mostly an automated system, but we simply don't let automation to work through exceptional situations, without human review. Therefore, lots of pilot training (and their presence) is still required.

But it's also hard to guess what exactly skill should the pilots have. It's clear that lots of things from manual aircraft do not translate well to automated.


Children of the Magenta Line seems very topical, especially wrt AI

>In 1997, American Airlines captain Warren VanderBurgh gave a lecture warning about the dangers of autopilot. He titled it "children of the magenta line," after the magenta coloured course line the flight computer draws across a cockpit display. He noticed that pilots were so used to following that line, to managing the automation rather than actually flying the plane, that they lost their own skills to manually fly the plane. This wasn’t a problem until it was a problem.

>VanderBurgh’s concern was not about autopilot itself but rather the complete dependency on autopilot at the expense of the internal knowledge of how it works and the loss of the judgment to know when to use it. In 2009, his warning came true in the very worst way possible over the Atlantic. When Air France 447’s sensors iced over, the autopilot handed a perfectly good airliner back to its crew, the pilots, confused and out of practice, stalled it into the sea.

-https://carlhendrick.substack.com/p/children-of-the-magenta-...

Children Of The Magenta Line -https://www.youtube.com/watch?v=5ESJH1NLMLs


Most people do not have the ability to understand a domain at the deepest possible level.

Just like kids eventually take over most of the daily tasks for their parents (for people in countries where you don't have any kind of retirement fund and assisted living, etc.). Same will happen here - on one hand because the human population keeps growing older without any replacement and on the other because all the new tech will be too complicated for humans to understand.

It seems unlikely that we'll be able to keep up with AI in any capacity. Your alternative seems like the inevitable outcome of AI succeeding, and it doesn't sound so bad.

Dunno man, the articles ai writes are empty cardboard crap, and the code it produces is hot garbage. “Cleaner upper of AI messes” might be a good future skill to have.

I don't think you can dismiss the output quality so easily.

Give yourself a few hours to write out some code in the old way. Can you really tell me that you could not have done a better job with AI in the same time?

Eg, implement chess or poker games.

Or write a component like a Swiss table.

It's not obvious to me that AI produces bad work at all.


That’s the thing - many times it’s not obvious. But go check that chunk of code that, sure, works, and you’ll find subtleties that just don’t look right. But hey - you do you if that’s the kind of quality you’re happy with.

And traditionally written code won't have those subtleties? Really?

As has been said before we've been cleaning up garbage code for generations and that doesn't look even close to changing anytime soon.

100%, does it matter if I cleanup the mess of another dev or of an AI? I don't think so.

Can other devs spew out crap at the same speed at AI? Cleaning up fifteen 10kloc pull requests per day? Yeah didn’t think so.

I hope people recognize this account doom posts what amounts to misanthropic propaganda at a very high rate. There’s never any question how the comment is going to be framed. It’s always, “give up little humans, your robot overlords are here.”

Yep, I've noticed that myself. It's getting pretty boring.

Yes, kids are the demographics known for long term thinking and rational decision making.

Where do you think they are putting their energy?

The answer is mostly "do nothing, the model will figure it out", with a side of "ask the model to check its work".

This matches my experience, where the job of the engineer is mostly copy pasting requirements, letting the model do the thinking, and then manually testing the results.



Yeah, about those I know, but what about cloudflare?

They hold your tls keys and can decrypt all your traffic. They're MITM as a service, by definition. They have to be able to in order to cache and forward appropriately.

Also to do DDoS mitigation. Being able to see the HTTP request, at least headers and path, greatly helps with distinguishing attackers from legitimate traffic.

It's a tragedy that there's no standard to allow partial decryption/nested encryption in HTTP, which would allow intermediate proxies like Cloudflare to e.g. only validate a first-level authentication token and rate-limit access to a given endpoint, but not decrypt the actual request body, backend authentication token, or response.

Also desperately missing: Authenticated static file caching (think: cdn.foo.com serves files authenticated/signed by foo.com). Subresource integrity only works for HTML use cases and is clearly not ergonomic enough to make a difference.


I'm not sure why the people in power would actually be making their own decisions; if the other country's leaders depend on superintelligence to make their realpolitik moves, you'd expect yours would need to do the same to match them.

So, the governance of a country comes down to AI alignment. You can't let the other guys get a leadership advantage through AI, or you'll start losing the competition.

When we get ASI, we're at the end of human decision making. Now's the time that we have to make sure ASI decision making is to our liking.


Why do people need to understand proofs? If Amazon improves package routing with new advances in graph theory, my cat doesn't need to understand it to benefit from better shipments of cat food.

Similarly, humans don't need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it's writing most of the posts on the front page of this website, and it's doing the bulk of programming at many startups.

We can't put this genie back in the bottle.


Well, for one, most math proofs don't have any practical applications, so a proof that no one reads is basically a digital paperweight. You might as well suggest AI write novels for other AI to read.

The hope is that some of them end up being useful; otherwise, nobody would be funding math departments. Mathematics typically anticipates and enables new physics and chemistry.

If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code, if it's not for any practical use?


If you're interested in the topic enough to comment on it, you'll probably find it worthwhile reading a mathematician's perspective. Here's the prolific Terry Tao: https://mathstodon.xyz/@tao/117219548485446992

They don't actually say anything about why anyone should fund this, though. I don't get why a society should worry about progress in mathematics if there's no practical benefit expected.

Maybe there's two kinds of math that we need? Useful math and navel gazing, and we can hand the first to the machines, and let hobbyists do the second in their free to entertain themselves?


It's hard to know what math is 'useful' a priori. That's always been the argument for supporting basic research. This is not why I am a mathematician however. I think there's intrinsic value into understanding something of depth and meaning, but the societal setup we have now that mostly agrees this is valuable is probably a very contingent phenomenon that is unlikely to last much longer.

Yes, so if there's useful math, you throw the LLM at it and use the results, no humans needed.

Humans can try to extract some ideas from the million line lean proofs, if they want to, I guess. But I can't imagine anyone really funding the human part of it.


It is the top tier of humans in these fields that are making the significant breakthroughs. The top tier of breakthroughs are not being post on here (which are nowadays usually short form articles of not incredible quality). The code at start ups is not commonly in the top tier of a breakthrough. AI can do averaged work and derivations off what has gone before which Maths works very well for as there is a clear set of rules. The same in physics if you ask AI for help adapting a simulation, yet it couldn't pluck the idea if no one has done it before.

Pure cope.

It's very clear that the AI still has no motivation beyond its prompts. Humans can mostly outsource their thinking today across a wide variety of topics, but they still need to express their desires.

Yeah, we're watching AI obsolete intellectual work in real time. Human brains simply won't be able to compete. We're not going to have to toil, we just ask for results.

This is no longer skilled labor. It's massively more productive when machines take over the thinking, but you need to look elsewhere for intellectual stimulation.


It seems like making progress on math is letting the AI run fully autonomously for a few days, occasionally asking it to keep going.

I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.


Are they actually autonomous? I’d say subject knowledge at the prompt stage plays a large part towards getting proper results

When Claude made progress on the Riemann conjecture, here are the kind of prompts used:

> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

And left it for a long time. Jarred isn't a mathematician, he's the maintainer of a janky JavaScript environment.

Here's the transcript: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...


Prompting for some of the results was almost the "Computer, do a breakthrough. Make no mistakes." meme. Just someone telling the model to keep trying a couple of times.

Unfortunately we don't actually know what kind of prompting was done for the more prominent results.


It’s definitely not how I work, I’d need to read the model responses and set a direction for the model to go

Same here, it's one thing if I'm just screwing around, but if I'm trying to do anything serious, I need to at least have a handle on what it's doing, and, when thinking traces are available, keeping track of any logical errors in the model's reasoning.

you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like:

7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.

Citation: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3..., section 6.2.2

You're not going to get a handle on what it's doing. The thinking traces are there to make you feel better about yourself.


Anthropic specifically calls out that example as

>"illegible reasoning in a few reinforcement-learning environments over long rollout"

Yet, I get the point that you're making: those tokens essentially are an internal scratchpad for the LLM which isn't required to logically lead to the output.

This video presentation of the paper you linked was interesting: https://www.youtube.com/watch?v=hUp3zh23aHw


I'm referring to models where the actual traces are available. Eg. I've been using a local Qwen3.8-Next-Flash lately.

Still not meaningful -- https://arxiv.org/pdf/2504.09762; even for local models, the reasoning traces are often filtered and summarized to sound sensible to humans. And even if not, they don't necessarily represent what the model is thinking.

Hmm interesting, thanks for the link, I'll have to give that a read.

Why is the model playing poker

That's how several major AI advancements have happened. I have seen no evidence that that is the fastest way to make progress right now. I expect that, much like chess engines, it will not take too long before AI is significantly better than AI + human. But right now, my bet is that we are still safely within the window where an AI + human mathematician team is still better than AI alone (at least for the case where the human has learned how to work effectively with the partner....something that this event could possible be good for teaching).

I suspect that the best progress will be made by a team that purely spends their time taking a list of open problems and promoting "solve <problem>", without actually trying to understand anything. Just keep as many problems in flight as you can across as many sessions as you can.

You can probably ask the AI to come up with a list of problems itself, and rank them by the likelihood of progress.


Then 20 years go by and you wake up one day with questions that you cannot get out of your mind: why did I start prompting the LLM for? Why did I need these random proofs for? What do i do with my repo with 2billion lines of Lean?

I think the purpose of an event like this would be to optimize the process so that it isn't just occasionally asking an AI to keep going.

That sounds like adding a bottleneck, unless you mean writing a harness that automatically asks the model to keep going, so that there's no humans involved at all?

What do you mean by invalidate? It means we no longer need to understand math, of course.

When Claude made progress on the Riemann conjecture, here are the kind of prompts used:

> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

It seems like the kind of prompts a high schooler could come up with. The prompter is the creator of bun.js, a kind of janky JavaScript execution environment. It seems like we no longer need to know the subject well at all.


What did the Medicis or Guggengheim's do other than be rich and idle at the right time and place in history?

The contemporary prompt is a port for patronage, not a point of inflection - with a similarly ludicrous symbiotic relationship defined by impression management on both sides. Same as it ever was.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: