Hacker Newsnew | past | comments | ask | show | jobs | submit | 6thbit's commentslogin

Buried in there, note you can opt to self-host your sandbox

https://developers.openai.com/api/docs/guides/agents-api/env...

That makes this much more enticing, and potentially eases transition between providers.


Then why tf do i need their api

The self-hosted environment is just the backend for shell calls the agent wants to make. The inference bits, thread persistence, and (optionally) mcp/tool calls happen from the API.

Some people/companies/whatever might like the convenience and scalability of managed solutions, especially if you're say, just building something simple like a Slack bot with your custom workplace tools/data.

Yes, the lock-in is real and only good for OpenAI, but there's absolutely demand for managed services where you defer the responsibility of security patching; scaling; uptime, etc to a third party provider. Just like why people use AWS/GCP/etc over bare metal in a colo.


Going to be GitHub self hosted runners all over again, you pay for the API and also pay for your own self hosting.

Native integration with their SDK.

For the lock in

If AI becomes a true commodity do you think OpenRouter/Stripe will evolve into this type of network that visa/mastercard represent today?

Any version with a bit more prose for a peasant like me to remotely pretend to grasp it

They (oAI) should just release all the prompts And internal reasoning for external audits.

They will, as soon as they get the same proof without it being obvious in the prompts they stole the research

So this is why forward secrecy matters eh

Running the authoritative dns for the zone seems elegant, although wouldn’t that imply you absolutely can’t use cloudflare or similar services to avoid ddos/bots?

Why have they not included ARC-Agi-3 on their index?

Clearly that would move things around.


Because playing a video game isn’t relevant to which AI people might want to use.

"Video game" is the medium, the challenge and test is figuring out how to win; when given no instructions, and specifically designed to be private.

ARC-AGI takes it very seriously: they've never tested Fable, because they won't run on the eval set without ZDR.

Another possible way to look at this is that any benchmark reporting Fable scores is potentially contaminated.


63 cents is quite cheap isn't it? If you compare vs Fable5.1 Max @ a whooping $3.30

Invoking simonw for pelicans pretty please.


The instant/no reasoning performed extremely well

    none 35.2%, $49,791 96.7%, $23,457
35.2% on the standard harness, that's above Opus 5 on high.

Since low scored much lower than none, and none scored ~ around medium, could none default to medium in the API? I don't think the new models can even have "instant" via API, unless they train them for that (there was one gpt5 variant called instant or something).

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: