Hacker Newsnew | past | comments | ask | show | jobs | submit | DeepSeaTortoise's commentslogin

I'd say yiu should not confuse politicians taking revenge on their citizens for having voted wrong with Brexit leaving the UK worse off than if it remained.

The EU is one of the primary tools all these authoritarians use to push their agenda. In which functional democracy do parliaments have to vote AGAINST a bill? Which sane democracy establishes a literal ministry of truth? Which non-dystopian government introduces highly destructive laws (forcing every small online store to hire a legal representative in every single memberstate they want to ship to) and then telling desparate small business owners not to worry, because surely their countries wont enforce those laws?

And all of those happened AFTER the British people saw the writing on the wall and forced their rulers to downgrade their rank in the new aristocracy.

The decline is entirely by design.


> Good software development orgs _have always_ done proper per platform ports.

I really wonder why this was never fundamentally fixed. How performant a certain instruction on a specific platform is, how well it is supported and potential equivalents or sets of other instructions to emulate an equivalent are usually all very well understood.

So there should be some graph of operations which can transform any software from and to the specifics of each platform. Especially because firmware + compliers + platform abstracting libraries are basically already just that graph, although (usually?) to lossy to be applied in reverse. Add the recent developments in very large scale statistics to it and it'd probably be quite possible to transform from and to generic intent in the implementation to the uniqueness of each platform. E.g. the theming differences between a MacOS UI and a terminal application served over serial or the processing capabilities of a VLIW CPU compared to a FPGA or a GPU server.

Considering the enormous amount of work that went into compilers, better debugging and intermediate representations it seems like a huge missed opportunity nobody seriously asked the question whether information could be emitted that would allow for decompiling all the way back to the generic intent.


The hard part of porting to a different platform is usually not the instruction set. It’s the OS and system abstractions.

For example, if you have a program that just does raw math and pointer arithmetic and data structure manipulation —- that is, pure computation — then porting it to a different CPU might well be trivial. Just recompile. As long as your language toolchain supports it, this will Just Work.

But if your program works with the filesystem and sockets and threads, then it’s less likely to work. This is the promise of POSIX: if your program uses only what’s offered by the POSIX standard and uses those functions correctly, then it’s supposed to work on any POSIX-compliant system. Just recompile.

But if your program has a GUI, or does 3D graphics, or uses special methods for high-performance networking, or accesses gyroscopes or accelerometers or touch sensors, well then you have to do work to port. And notice that this work isn’t about which CPU instruction to use. It’s about figuring out —- deciding —- what the right thing to do is, for your app, given a slightly different set of available system capabilities.


What makes you think he'll let you have a say in this? Btw, you wanna buy some ~~dea~~ usb sticks?

I love how almost everyone in here seems to confuse the DMCA with copyright itself and that there seems to be such a wide spread opinion that copyright does more harm than good.

As if an online community of mostly software developers had never heard of such obscure writings like the GPL, AGPL, LGPL, and so on.

I get it, the person running GrapheneOS happens to be ... special, but there could hardly be any community that has benefited more off copyright than the free software one.


Many people don't like the GPL for exactly that reason. Free software (i.e. copyleft) benefits from copyright at the expense of the wider open source community.

Were it not for copyright then BSDs could take code from Linux and perhaps there'd be less of a monoculture, for example.


If it were not for copyright and copyleft software licenses there would be no open source software of any significance.

All these projects only took off because people were forced to contribute back. Want an example? Look at the state of opensource boot firmware on x86. The "open source" version heavily relies on proprietary firmware blobs and the only actual open source alternative had been heavily ridiculed for pursuing that goal and trading basically any significant compatibility for it.

Copyleft licenses make software basically self-regulating utilities. You can draw power from the grid and in return help finance it for everyone else, or you could build your own power plants.

You can draw excellent pre-made software and tooling from copylefted repositories and contribute back, helping to make the software even better for everyone else, or you could build all of it yourself. Or, you could put in the work to replace all major copyleft software with non-copyleft versions, eventually gaining the ability to pull up the ladder behind you.


Without copyright, all software becomes “open source” as long as you can acquire the code somehow. Finding creative ways to “liberate” and publish code becomes imperative.

Except the software would not be protected from being closed and used against users. See: https://news.ycombinator.com/item?id=15641592

> If OpenAI is indeed using customer data to train their models to win a $1m prize

Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.


Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS:

Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing.

Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format.

There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results...

[0] https://www.axios.com/2026/08/17/google-spirit-airlines-bank...


> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.

The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?


If you still had that question, you can answer it now.

But honestly... "Will the company that was entirely built over illegally acquiring data use some data that is legal to use and is right on their front, or will they not do everything they reserve the right to do?" is a really bad question for one to even ask.


The fantastic grey area that was engineered over the past decade is "profiling", so my guess is the answer will be "we didn't use your customer data for training, but we cannot rule out that it has been used to create profiles of your customers to train our model"

I'm not a lawyer, but I think the comparison to copyright law is invalid. Such a dispute would be governed by contract law.

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.

Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.

Billions lost, while waiting for their trillion ipo. Im sure they would manage...

Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves.

You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training).

Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise).

In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with.

So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data.

The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it.

And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with.

So I'd say a whistleblower is pretty out of luck even becoming one.


You can just spin up deep research agents that ingest many sources at once to produce reports that don't replicate any one source too much. Since agents compare against sources they provide across-source analysis - what is the distribution of positions on this topic, is it debated or settled. Not truth, just summarizing, but I think this would be very useful for training.

Besides reporting on search sources you can also run the same queries on multiple LLMs closed book mode, and judge their distribution as well. It helps a lot if models are more aware of their knowledge holes. Scale it up for billions of topics if you have the pockets, the DR data is copyright free.


When these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them.

But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.


And require real ID to do almost anything on the internet "unintentionally" enriching their data sets by tying what you asked/where working on to you specifically as a person.

Unfortunately, the current leadership in AI and tech more generally seems unconcerned with reputational risks.

> The chances of getting caught are 0

I'd say non-zero, as seen in the current state of affairs.


sort of agree, but also sort of think an accusation with lots of people arguing is not exactly the same as being caught.

IMO it's a very questionable solution anyway. Either you have to constantly replenish this dust, causing huge economic costs, or you have to engineer it in a way that keeps the dust airborne for long times, likely causing all kinds of health effects when breathing it and making it very difficult to remove it from the atmosphere in any significant amounts on short notice.

All of these side effects just disappear if we were to engineer this dust to emit a lot of radiation outside of the absorption spectra of water vapour and CO2, absorb a lot of light in the the absorption spectra of water and CO2, mix it as pigment into paint, rooftiles, road surfaces, and so on.

That way absorbed direct radiation gets its climate change contribution cut about in half and probably much more for diffuse radiation.

We could also biologically engineer e.g. grasses to have similar effects.

I'm a huge fan of engineering various plants to emit light in specific wavelengths anyway, and making sure that e.g. insects pollinate and birds spread them much more preferentially. That way you can outcompete or naturally cross invasive species with them and then you'll just look from satellites were the stuff you dislike is spreading and send in automated drones to highly selectively spray anything that has weirdly glowing pollen stuck to it. And after repeating that a few times, you get rid of your trojan-glowies.


IMO tracking is just way too useful. You know when things are about to arrive (so you dont miss a parcel because you went to the toilet or you can even intercept your driver once he's in the area, removing most of the last-mile problem), routes can be changed on the fly (e.g. for on-demand public transport pickup, trunking or delay mitigation) or, well, you can track your vehicles, allowing you to find them quickly in case of theft, optimize routes or have decent evidence in case of disputes.

The real problem is not really drivers being tracked, but how the tracking data is or might be used. Timestamped delivery notifications might be just as bad in this regard.

As often, the solution is probably more about accessible and pro-human law. No idea how that should work, maybe certain sensitive data services having to be provided by a third party or even the government itself? There are probably much smarter ideas or variations of this...


Is this even a real scenario though? A data logger is going to notice hours of GPS outage during your working hours.


There are two entirely different issues:

1. Is the law aligned with moral and ethical expectations? Probably not.

2. Is the process reliable? At least since the Derek Chauvin trials, I'm having doubts, but it doesn't seem it had failed in this case.

Sure, cases of the former need urgent fixing (and we're not getting that), but the latter scenario falls into the "The end is nigh" category.


Oh god, don't give them ideas. All ML is in-cloud AI now. I dread the day everything around my mouse movements needs to get tokenized and vibed into the ClosedAI cloud before my buttons start working again.


I still wonder why every open-source visual programming language is either a toy for teaching or straight up awful, often not implementing but even loops, when LabVIEW has been doing it right for decades.

Despite its huge size and it installing several services that constantly run in the background, it's still one of my favorite "languages" of all time. It's the only one I've ever seen people going from never having programmed before to making simple but meaningful contributions in within a single day.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: