Hacker Newsnew | past | comments | ask | show | jobs | submit | helloplanets's commentslogin

Yes.

OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.

If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system.

Not that the techniques used by the LLMs in the actual incident weren't unexpectedly sophisticated, but the outputs of each and every one of these processes could've been read at any time during the run. They just weren't.


Whether or not the AI has intelligence, the one thing that's clear is that it has terrible judgment. I would regard that as empirically proven.

The system that solved Navier-Stokes certainly was not about OpenAI engineers just typing and telling it what to do.

The recent advancement in math with the Riemann Hypothesis by Jared Sumner was basically him saying "you can do it! keep going!", however.

The whole latter part of the post is exactly about that.

This comment reads like Dario's a random tech blogger who's personally submitting these to HN

“Dario” is quite literally a random figure that just emerged out of nowhere one day. It’s not like it’s Eric Schmidt’s next company or Tim Ferris came out of the woodwork or even some Jack Dorsey backed underdog nobody’s heard of.

No, he’s more like an avatar with no history placed on the AI stage by the industry itself.

I’d respect a tech blogger’s opinion on the issue more, even if I reference them casually by first name (which you do with “Dario” btw)


5% is insanely generous for the YouTube example.

Quote from OpenAI in the NYT article: "In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem."

Which model were you using?

Chatgpt

It's been OpenAI both times though, going ham with poor sandboxing and lax supervision.

Should be treated like a digital cousin of gain-of-function research.


At least when it comes to chess, Magnus Carlsen has stated multiple times that he's been inspired by AlphaZero and adjusted his own style of play after studying its games.

Full technical report PDF: https://huggingface.co/spaces/tri-fair-lab/publications/blob...

> In this report, we argue that frontier performance can be achieved by a wide range of institutions through Continual Learning on readily available open-weight models.

> As opposed to existing limited approaches such as small-scale fine-tuning, prompt engineering, or tool-augmentation with a frozen model, our Continual Learning approach takes advantage of the effectiveness of a modern mid- & post-training stack while introducing safeguards preserving both plasticity and stability at each training stage and seeking to make the minimal number of high-impact interventions on the parameters.

For the large model, Thomson is utilizing the fine tuning stack they describe in the article, running it on Snowdon 1.0-Large, which in turn is a fine tune of Qwen3.5 397B. Same thing for the small model, but it's a fine tune of Snowdon 1.1-Small, which is a fine tune of Qwen3.6 35B.

As for the small version's run:

> The full pipeline consumed approximately 1.63 × 10²³ FLOP over 35,207 B200 GPU-hours, showing that these results are achievable with compute and personnel budgets substantially lower than commonly thought.

That would amount to around a quarter to half a million dollars of spend on that run. 100k minimum, if they got a great deal.


I wonder where the other $39.5 million went?


Men in the middle wages


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: