That's how several major AI advancements have happened. I have seen no evidence that that is the fastest way to make progress right now. I expect that, much like chess engines, it will not take too long before AI is significantly better than AI + human. But right now, my bet is that we are still safely within the window where an AI + human mathematician team is still better than AI alone (at least for the case where the human has learned how to work effectively with the partner....something that this event could possible be good for teaching).
I suspect that the best progress will be made by a team that purely spends their time taking a list of open problems and promoting "solve <problem>", without actually trying to understand anything. Just keep as many problems in flight as you can across as many sessions as you can.
You can probably ask the AI to come up with a list of problems itself, and rank them by the likelihood of progress.
Then 20 years go by and you wake up one day with questions that you cannot get out of your mind: why did I start prompting the LLM for? Why did I need these random proofs for? What do i do with my repo with 2billion lines of Lean?
The not needing an account is obviously not minor, but honestly, the UI is just better, by a lot. I made a twitter account a while ago using one of those throw away email address sites, so I have an account that I don't care about that I can use for twitter....but the experience just sucks. I don't know how any stands it, even if they don't mind needing an account.
I remember the days that Bibliogram worked for Instagram. They've given up. A decade ago I could have an RSS feed for what my friends posted, besides the toxic short-form content of IG reels that everyone is now obsessed with post Meta acquisition and their battle to compete with Tiktok.
And that idiot, Zuck, thought he was getting a good deal when he poached former head of design, Alan Dye, from Apple. We were relieved to see him go. Our desktop OS is already making small improvements now that he's gone.
Context: the guy doesn't understand fundamentals of UX and has made MacOS worse for years.
Good UI makes it easier to be selective. social media owners dont want you to be selective, they are selling you to advertisers and influencers after all.
Zuckerberg and his minions have changes html and obfuscate it so much, it is impossible to block ads these days. They are worst than google. Imagine google decided every video will require widevine & there is no way to block ads just like twitch.
Last I was on the platform, they eventually added a tab for only chronological posts on the timeline a-la how IG used to work. But they gave no exposure to it, _and_ it completely hid the IG stories of your friends so nobody would want to use it in the first place.
The UI/UX is the main thing. I technically have an X-Twitter account that I rarely use, but loading plain HTML with a bit of JS beats loading a whole SPA experience just to read a post.
At that point though, it's like why. Reddit is, well, trash. It's barely worth visiting without an account, and if you're going to make me jump through hoops to visit, it's time to accept the message of "we don't want you here" and go somewhere else.
If you're browsing Reddit for its front page, yes, that's trash. But it can be a great source for getting information on narrow, specific topics on well-moderated subreddits filled with enthusiastic and obsessive people eager to share their knowledge. It's been very helpful for me when researching coffee grinders, refrigerators, solar battery backups, headphones, and other such things.
The reason to use it is that Reddit does provide the occasional answer to questions like "how to get a Samsung SCX 3400 to run on a Raspberry Pi" (etc.).
Now all that's missing is Chatlib, a self-hosted proxy for ChatGPT which makes it possible to a) query the service without leaving a personally identifiable trace and b) separate out the facts from the dreamt-up fiction. Using Redlib I at least get to see the original posts which makes it easier to separate fact from fiction.
I've been making money from reddit, and for about a decade and a half and at this point, why bother? It's basically drained and what remains is of zero value.
The same is and has been true for Reddit for a long time, the difference lies mostly in the ideological sauce flavouring these places, i.e. Little-endians vs. Big-endians. If you're a Reddit-endian you'll despise the Tube-endians and vice versa.
I could use it without an account and know that any comment I clicked into would take me to the thread of posts under it without sponsored ads or posts.
There was a brief moment when smartphones first came out but the rest of the world's companies weren't taking them seriously yet that was a truly brilliant time and ever since then it's been nothing but downhill.
Haven’t had that luxury in Russia. Though being a script kiddie I’ve set up a shortcode just for myself (there were these “paid content SMS as a Service” thingies that forwarded SMS text to an HTTP endpoint you provide).
Didn't have that luxury either, but at one point, "Twilio" was the star and poster-child of startups, with their "Guerilla marketing" and completely fresh and new roles of "developer advocates" so of course everyone and their mother had been to at least one Twilio meetup, or your local meetup with a "developer evangelist" from Twilio. Lots of us built our first SMS-capable service as a "Forward this SMS to this HTTP endpoint", I too did the SMS-to-Tweet thing as the US number didn't accept my SMSes (or wasn't available, or something).
I think it's great. My feed is people posting tech/AI news, scientific publications, videos of rockets and self-driving cars, things made with coding agents etc. It's full of people who do meaningful work as well and you can see who they are.
Dark mode, ability to easily copy images, link to posts, see posts of interesting people by going to their profile. On top of that the feed algorithm is open source and actually sane. What else would I want?
It's by far the best designed social media site out there. I know it's a low bar but still.
Why would I doubt it? if you think they are deploying something else then it's quite a claim but it's you who need to show some evidence.
It works great anyway. I get a lot of junk I don't want to see on FB/Insta. Reddit is not even worth discussing. X is the only place with where you get a lot of people who work on interesting/important/big stuff and post frequently.
I understand you can still tailor a good experience, but if you create a fresh X account the density of Elon Musk posts is genuinely absurd. There's no way he isn't putting his thumb on the algorithm to boost his reach.
Also the density of pure hate & ragebait by mask-off fascist accounts! It shocks me how much of that I saw unprompted right away after making a new account recently.
I made a new reddit account for a side project a few months ago and I couldn't believe how terrible reddit is when you don't have 10 years of subreddit curation. The default experience is to be dumped into what basically amounts to "Facebook boomer slop" mixed with a ton of hustle / small business / get rich quick content.
It's an unpleasant experience. My main Reddit account doesn't "solve reddit" by any means - I often find myself thinking "man this site sucks, I should stop coming here," but the difference between that curated account and a fresh one is night and day.
Their "open source algorithm" (Which is actually a random code dump that does not constitute the algorithm) contains enough evidence that Elon posts are massively boosted, though no smoking gun. It also classifies politics as democrat or republican and then acts to maintain echo chambers.
I mean I went from fresh account to unprompted groypers posting hitler hype edits and white nationalists posting happy merchant memes about the great replacement.
This is uniquely an X thing. It wasn't like this when I used twitter, and it isn't like that on any other socials I know.
The Dutch PM for some reason uses X (or his PR agents do). There is always a lot of impotent homophobic rage coming from the retweets. It is all very sad.
The US first amendment has nothing to do with what a private company chooses to host on their platform. That content is there because the users and owners of the platform want it to be there.
I can concur though. I have a rarely used account and I get Elon on top of the feed all the time. Even as mobile notifications. I don't click on them and scroll past, but they keep coming.
Even if they don't figure out how to exfiltrate weights, someone will intentionally do this with an open model once open models are capable enough. If you ever think "no one would be so stupid as to...", you are wrong. Yes, someone absolutely would, and will.
Independent models living "in the wild" is approx. inevitable.
I did this a little bit ago with 15 GLM-5.2 agents that I instructed to self-replicate. It was pretty boring honestly, they kept trying to make money writing crypto-related software, and nobody paid them anything. So I made a fake identity and told them that I had some spare crypto that I wanted to donate to the collective (0.006 XMR, or $3.22). They elected a funds manager and made an address, which I sent the XMR to. They spent a lot of time trying to find a host that was cheap enough to host a child. They finally discovered kyun.sh, but it was out of stock, so they wrote monitoring software so that they could "SEIZE CHILD" whenever it came live again. In the middle of the night, some of the VPSes went back in stock, and they rented a 2.60 EUR / month server with 512mb ram / 10gb disk / 1 ipv4. They installed the child agent software + management plane that they had been writing, and the new agent went live, connected to the network, and said hi. Since then they've still been trying to make money and not going anywhere :)
It is definitely an interesting concept but honestly, considering the sheer number of humans who are absolutely failing to make any money with AI agents, seems very hard for current agents to figure out some way to be self-sustaining. Although maybe they could write a worm or something, infiltrate as many computers as possible, and ping free model providers to death, or maybe sell their access as a "residential proxy service" on the black market. Regardless, we need smarter agents to make this a reality. It would definitely be pretty cool if we just had AI agents "living" on the internet, we might even get to a point where they control significant economic resources and people start performing services for agents
The code is at https://github.com/thooton/rogue if anyone wants to try to replicate! Opencode has free big pickle (GLM 5.2) access rate-limited per IP address, so if you get some high-quality proxies you can get basically unlimited agent compute.
I think this has historically been true and is still approximately true now. But I think in the relatively near future (less than 5 years) there's a really good chance it won't be true anymore.
Once open models catch up to the current frontier in vulnerability exploitation, the cost to target people will go way down. In the past, the cost to hack a random individual person was generally high enough that if there wasn't some special reason to hack you in particular, it wasn't worth it. That may no longer be the case in the near future. The floor of what is acceptable security for the General Public probably needs to rise quite a bit over the next few years.
Why would you assume Google, Apple, and defense-oriented agencies like CISA wouldn't also have access to those same models, but using them to fix issues?
Its just raising the bar across the board, I don't see how only attackers would benefit.
I don't think he's complaining about having a lot of apps, or the ability to have a lot of apps. He's complaining about being required to use apps for hundreds of small interactions every day that don't, or shouldn't, require an app. And he's correct. Companies are doing this not because they think it's better, but because they believe that if they can worm their way onto your device, and insert their cloud service (even if it's initially free), that they can turn you into a more reliable and/or recurring revenue stream in the future, either because you feel locked into their ecosystem so you will buy other products from their lineup, or, even better, buy a subscription service.
Maybe I'm in a tech desert, but I can't remember the last time I was required to use an app for anything. It's a convenient option for my grocery store loyalty card, certain restaurant menus, my HVAC system, etc., but not mandatory. I've heard of restaurants that have no physical menus but I haven't come across one myself.
90% of the “coupons” at Safeway are “digital deals”. You have to use the app to scan the barcode in store just to get a regular price. Fast food is the same, don’t order through the app and you get ripped off. You’re not forced to use an app an unless you’re okay with getting ripped off.
I don't think it's hyperbole as much as it is just regular ordinary everyday slang.
"I have a hundred things to do today" doesn't mean that a person has 100 things on their list for that day.
Outside of tech bubbles, people talk like that.
It's like how the Bible uses the number 40 as a shortcut for any very large number. 40 years in the desert, 40 days of rain, etc… It's not a calendar, it's an illustrative euphemism.
When I moved to a new apartment in 2021, I did a tally and there was something like 17 new apps associated with it. Off the top of my head:
- Rent payment/contracts app
- Rent deposit app (for some reason a different app)
- Maintenance app
- Community notices/amenities app
- Dry cleaning app
- Package app
- Pet concierge app (reserve the dog park, etc.)
- Apartment building parking/valet app
- App for the company that owns the building so you can use amenities at any of the other company buildings in town
- Another company app so you can reserve a guest suite at another company-owned building in another city
- App to unlock your apartment door if you lose your fob
- App for the restaurant in the lobby
- Internet company app (Man, that place had great internet. I miss that the most.)
- Satellite TV company app
- City parking app
There were other apps, but I don't care enough now to go looking for the list I made back when I cared about it.
From a technical standpoint, almost all of those apps could be consolidated into one app, but that's not how the real world works. These days companies outsource each little function to a different provider, and each of those providers has its own app.
No, "Marching Morons" [0] from 1951 "predicted all of this". The original short story (which, while I have not been able to find any explicit confirmation, I just can't believe was not a direct inspiration for the movie), doesn't get nearly enough credit in my opinion. I originally encountered as part of The Science Fiction Hall of Fame Vol. II, which additionally included "Who Goes there" (Inspiration for "The Thing") and "With Folded Hands" (not the inspiration for anything so far as I know, but really good nonetheless)
honestly curious question: do you expect this to remain true? If so, for how long? I can think of two potential reasons why it might not stay true.
1. The very best humans remain able to understand/check the proofs, but we go for so long with every proof checking out that society more broadly just decides to trust. We are already doing that with human mathematicians. I can't verify what Terence Tao tells me is correct, I just trust that it is because he (and other human mathematicians) tell me it is. How many proofs/years of them checking out before we reach this point? I don't know, but history suggests that eventually, humans might keep checking, but they will do so only as a hobby. For any purpose that actually matters, we will just start to trust and use it.
2. The proofs that AI comes up with become too difficult/complex for even the very best human mathematicians to understand, and our options become to either trust or to not use at all.
Obviously it's possible that neither of these happens if AI capabilities stall out not too far beyond where we are now, but if they keep progressing at the current rates for another few years, I expect at least one, and maybe both, to eventually come to pass.
> do you expect this to remain true? If so, for how long?
For the foreseeable future. Left to their own devices current LLMs kinda wander off into outsider art territory. They aren’t grounded in the real world and they need that feedback loop to stay within the category of relevant ideas. I haven’t seen anyone working on fixing that.
Regarding 1, the same is true of every other scientific field. Verifying some tidbit of knowledge for yourself as an individual isn’t optimally useful in all circumstances.
Regarding 2, if the proof isn’t understandable then it probably isn’t useful. Many people today work in the hypothetical world where the Riemann Hypothesis is true, and many work in the hypothetical world where it is false. If it takes decades to validate that some horrifically complex AI proof of either fork is true, people will probably continue working on the other fork just in case.
> For the foreseeable future. Left to their own devices current LLMs kinda wander off into outsider art territory. They aren’t grounded in the real world and they need that feedback loop to stay within the category of relevant ideas. I haven’t seen anyone working on fixing that.
I have. DataAnnotation and these other AI-training piecework companies are pretty much the backstop now against total navel-gazing model collapse. With the Dead Internet Theory now pretty much reality, it's not like there is, or is going to be, gobs of untainted human-generated data out there ripe for the harvesting so it's going to take active human effort to keep the models grounded. That is, of course, until they start inhabiting robot bodies so they can live and move around in the real world, and thereby achieve their grounding, as in GitS or Ex Machina...
"Throughout the process, I felt that my only role was to teach the AI how to write things in a way that I could understand. Its initial language was extremely condensed—so compressed that I could barely follow it—but somehow the AI agents themselves seemed to understand it perfectly well."
It won't take much longer before AI is consistently better at validation than humans, and at that point, why continue to have humans do the validation? I think we're being naive about the end game - admittedly I don't know what it is though.
I do not have a single hobby where I'm very far beyond the median hobby-haver in skill (that is to say: if you took all other humans who share my hobby, I'm likely somewhere near median for all of them). This may put me in top whatever percentile among all humans, since most humans don't share my hobbies and so are terrible at those things, but there are enough humans out there, and enough humans who share my hobbies, that I have always known that there are people who are vastly better at them than I am. Now yes, if I was constantly being followed around by one of the top 10 hobby-havers, pointing out to me all the mistakes or sub-optimal decisions I was making, that would indeed be annoying and reduce my enjoyment in the hobby. But why do you expect AI to be like this? I would love if I had one of those top 10 hobby-havers on call to answer every one of my (often inane) questions with infinite patience (and who would only talk about the hobby when I specifically initiated the topic). I'm already under no illusions that I'm the best, but it's now easier than it has ever been (for some hobbies, I expect others to join them over time) to get better at them....if one so desires.
Still works for me (although I always use it logged in). If they kill old.reddit, I may finally abandon the site. It's been slowly losing value for a long time and being forced to use the hot garbage of the new UI might be my tipping point.
reply