Hacker Newsnew | past | comments | ask | show | jobs | submit | irthomasthomas's commentslogin

It took them two years to finally get him out with the help of the government in Shenzhen

https://www.nme.com/news/arm-china-finally-ousts-rogue-ceo-t...


And this is why building governance out of a pack of cards is a risky business...

If you cannot escalate a governance issue to a higher power, the controls are not worth the paper they are written on.


Wait China pulled off a coup inside ARM?

> For instance the task naming in the task file starts with an optimistic 1, 2, 3, 5, 5a but then eventually gets to 8a, 8a1, and then ends up with 8b2c2b3 and “8b2c2b2b checkpoint1”. The code that it produced got ever more wild. I don’t want to bore you with what it tried to build, but here are some example pieces of the interpreter changes:

  Hardcoded constants everywhere
  Multiple same-line macro invocations in C
  Random indexes in production code
  Hideous tokenizer code in C
https://lucumr.pocoo.org/2026/9/7/astra-why/

Astra scores the same on DeepSWE 1.1 (~75%) as Gemini Flash 3.8 and Deeepseek Flash 4.1 So general coding ability has plateaued, for now. Also consider the context windows. 1M token models where a breakthrough two years ago. Today they are still limited to 1M. In fact, if you don't want intelligence to drop off a cliff, you are really limited to 200k tokens.

Gemini Flash is a joke for coding. If you can get the same output as you can get with Sol/Astra I'm impressed. Not to mention that Antigravity is awful.

It is not a universal opinion at all that general coding ability has plateaued.


Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months

two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?


Just knowing that there had been progress is enough to have an idea that throwing more compute at it might work (OpenAI had previously tried all the Millennium Prize problems with somewhat limited compute and failed).

It's comparable to Magnus Carlson saying that if he wanted to cheat, all he would need would be for someone to tell him to spend more time thinking about a specific move (just a wink would be enough) as an indication that a computer had found something interesting.

It's as-if after OpenAI first failing on Navier-Stokes (which OpenAI had just tweeted about 2 days earlier!), someone winked at them and said "you might want to try a little harder ...".


When reading human comments, we should be generous; when we read corporate texts, we may assume paltering.

(TIL: paltering: exact and technically correct statement usage to create misleading impression)


The comment you replied to quoted "no user inputs after July 3rd" with no restriction to Buckmaster or Codex.

Obviously the result of OpenAI's investigation was that no usage data has interacted with the system after that date.

What else do you expect them to investigate?

If Buckmaster and co. provide their chats, OpenAI could potentially search for them in the anonymized opted-in usage data. Then they could say if any data has been used.

By all accounts individual usage data does not have the direct impact on the model most here fantasize about. To prove this, OpenAI would need to do new training runs to replicate the system used minus the particular usage data in question, if it exists, and then benchmark this on the problem again.

Potentially multiple times, in order to reach a conclusion.

The cost might be in the hundreds of millions.


Openai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts...

hmm I'm hoping there is a bug on their API because my first impression is not good. I asked it to return bash code between <bash></bash> tags. It is failing frequently and writing it's own tool calling format instead.

Quite a flex calling their GPT-6 competitor "Flash"! But it is faster than their last flash model due to a combination of architectural innovations including engrams and a new encoder/decoder design that uses 8B parameters for prefill and 16B for generation.

This is definitely not on par with GPT-6 astra. Not with GPT-5.6 sol either. But probably will set as a new baseline for modern API based LLM because it's so cheap.

Not on par, but in the same league. Astra is way ahead on visual tasks, but scores the same as gemini and deepseek on DeepSWE.

It's bench-marking near sol

Has the method for extracting the COT been blocked, now? Otherwise why could we not generate some fresh samples?

I'm not sure how viable it still is. Perhaps it's still possible, and perhas that's exacly what they did in wtich case my objection falls, but I don't know.

Chutes.ai models are served from a Trusted Execution Environment, so the GPU owners can't see your prompts.

And deliberate or not it is still plagiarism by the sound of it.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: