Hacker Newsnew | past | comments | ask | show | jobs | submit | otabdeveloper4's commentslogin

> you're just not using the latest model, bro

Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.

P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.


> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.

The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.


The tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price.

It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.


I think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.

I would run into the use limits very quickly, and (for Anthropic) have to switch frameworks.

By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.


No argument that DeepSeek is much cheaper, but you certainly get what you pay for.

artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.

Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.


It's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...

Where did he claim otherwise?

Take your meds.


Impressive how feverishly you defend the steaming pile of shit that Gemini is.

> everything you believe is wrong, here's our god, you must follow and believe in him or you're going to hell

"Democracy" and "human rights" is basically that but on steroids. Surely you're not a cultural relativist that simply going to allow uncontacted people to be racist and misogynist willy-nilly?


Those things aren't at all equivalent, so I'm not sure I understand your point.

The very first thing Garage did after I installed it on a test three-node cluster is get corrupted and lose files.

No thanks.


> Fact: There was no world war.

Akshually there is, and WW3 has been going on since 2010. (Mostly in places that aren't Europe.)


Some people believe WW2 started many years earlier, but most historians don't put the start date at the conflicts going on before the war became intercontinental with sides working together on a global scale.

War has still been going on since 1945, much earlier if not always, so akshually WW2 never ended or is just how it's always been? (mostly in places that are not "first world" / The West)


There's a simple and very logical definition of what a "world war" is - it's when international law completely breaks down, and the only way to get back to a semblance of order in international relations is to have the victors of the ensuing shitstorm impose a new set of rules in some sort of grand finale pact.

1945 was clearly that. 1920 with the League of Nations attempted to be that, but failed.

Whatever happens after the current turbulence we can be sure that there won't be a UN and NATO and a World Bank at the end.

P.S. The first modern world war was the Thirty Year's War.


> Whatever happens after the current turbulence we can be sure

Have you put your money behind this on future gambling sites? How can you be so sure?


Because the UN, NATO, etc., are clearly not working now. They're not gonna be fixed to work as designed because that's quite impossible in today's world.

> claims to be

Yes, quite the smoking gun of a load-bearing phrase.


Sure, but you can check yourself if you're worried about it.

All vibe-coded software is complete unusable trash. There are no exceptions.

So no, we're not advancing, we're in fact regressing. The crash and disillusionment of the post-vibe-code world will be immense.


If you ask an AI if "P=NP", it will certainly give an answer. It has to, because the answer is either "yes" or "no".

If you ask it to then prove its assertion, it will then certainly output the tokens that look like a plausible proof. It has to, as it is programmed too.

This "proof" might even be thousands of pages of very technical looking and professional sounding jargon.

It's not a real proof though, and as an artifact it is 100% useless to both the field of mathematics and humanity.


That's why this proof and all the other high profile AI proofs were written in lean, which allows you to verify them in a formal deterministic language. You still might not understand it as a human, but any computer with a simple processor can verify the proof's correctness. And from there you will undoubtedly see other people make sense of the proof's key steps using AI too. And the models might even pick up on further details useable for other proofs that humans didn't see. In the end I'm 100% convinced that abstract math will eventually be primarily done by computers, similar to how linear algebra and numerics have been done exclusively done by computers for a while. Noone would even consider multiplying a 100x100 matrix by hand anymore, if only because it is much more likely that you as a human will make a mistake.

That's assuming that "formal deterministic language" is equivalent to mathematical reality, or that it even reflects it correctly. It's the accepted view nowadays but akshually quite the hot take.

> The elephant in the room

...is the fact that the people pontificating on AI and math the most are also the people who understand AI and math the least.


No, Google will just do "product placement" in the AI output.

P.S. Google's real problem is that their web search is absolutely atrocious and completely broken. Even Qwen is better at searching the web than Google's AI search box.


If people are actually dumb enough to believe an LLM chatbot is god then it becomes a self-fulfilling prophesy.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: