How much of the “danger” is from better models vs the harness?
Isn’t the current risk due to how AI is configured, like giving it a full set of tools and internet access and a goal to hack stuff?
If we think we need laws or gate keeping, why isn’t it at this level? I already can’t ddos someone or fuzz their server or whatever right, I imagine if I threw equivalent compute at old school hacking I’d just get arrested.
The quality of the “frontier” model doesn’t really matter, they just generate transcripts, they can take no action.
If this was real they’d be calling on people to stop hooking them in to “dangerous” harnesses as opposed to pausing research. But it’s not.
What in this proposal gives you the idea that they're only talking about pausing research itself? It doesn't say that at all.
But yeah, directionally I think you're right. It's insane we ever let these things connect to the Internet or to write, never mind execute code.
But for the economic counterargument, the distinction doesn't matter much. Continuing research but constraining harnesses is just accepting nearly all of the costs and foregoing nearly all of the benefits.
Wouldn’t it most likely mean their progress is slowing and they want to make it appear as a deliberate decision for the good of humanity instead of admitting growth follows a logistic curve in advance of their IPOs?
Interesting. But even Elon agrees now. Is it really a commercial intent or something deeper? Only non-speculative answer can come from someone inside these labs with nothing to gain or lose.
Closes person I can think of is Andrej (he agreed too).
Is there any indication stretching leads to improved health outcomes like resistance training does per the article?
Everything I’ve seen indicates it’s a “nice to have” if one has unlimited time but if you’ve only got an hour a day or whatever for a workout you’re better off spending it all on resistance training.
The article is about muscle hypertrophy, not health outcomes.
>if you’ve only got an hour a day or whatever for a workout you’re better off spending it all on resistance training.
The health benefits of resistance training plateau around 90-120 minutes per week. And don't forget about cardio, which is arguably more important for reducing disease risk. I would guess that maintaining flexibility helps reduce risk of injury and mortality from falling, which becomes increasingly likely as you get older. But there's not much evidence that increasing flexibility beyond a normal functioning range has any ancillary benefits.
I work out at home with dumbells, the one thing I miss vs a gym is some way to do lat pull downs. I can’t do a chin up and don’t feel like I’ve got a good way to work up to it.
I also don’t work legs super seriously, I have 2x75lb dumbells I use for squats and deadlifts, I could do more if I had heavier weights but that’s not a priority.
For progressing to a chin-up or pullup, have you tried negatives? Or rubber band assist? Or, if you are willing to get rings, you can lower them so some of your weight is supported by your feet on the ground (body straight like a plank), changing the height of the rings will increase or decrease effort.
The modern way to progress to chin ups normally involves bands and eccentrics, but you can also just do partial reps (starting with the top of the movement) and gradually work on adding depth.
(This also avoids the stress on your elbows that can happen from the bottom of the rep if not controlled well.)
If you're quite far off starting with bodyweight inverted rows and working up in reps will train your arms and elbows in a very similar way from an easier position.
Anything "partial" gets a bad rep from people ego lifting with weights but it can be a legitimate way to work up to calisthenics exercises and saves fiddling around with bands
If you haven't tried single leg movements, that's the way to increase load on your legs when you don't have more weight available.
Moving fast is also a good way to increase intensity without increasing weight (you can clean less than you can deadlift, for example). The training effect is slightly different, but probably more useful for the average person tbh.
Chin up progression is eccentrics and band assists. They'll get you to your first chin up faster than lat pulldowns since you're actually doing the movement.
Erlich Bachman: Jian-Yang, what're you doing? This is Palo Alto. People are lunatics about smoking here. We don't enjoy all the freedoms that you have in China.
Anyway, when the talk of age verification first came out, I moved off of Claude and use either local or Chinese frontier models via open router (which probably isn’t safe for long now that they’ve been acquired). I’d almost forgotten why, this is a good reminder.
With HF bought as well, I wouldn’t be surprised to see them try to gate open models by age too. Going to have to sign up for modelscope
When I saw the 192GB I first thought they might have beat Framework to market with a Gorgon Halo machine with unified RAM that Framework has teased. This is just a GPU workstation.
Funny how they include nvidia, micron, and AMD revenue as “AI revenue” and that it represents that majority of industry revenue but presumably a big chunk of everyone else’s spend. Almost might as well include electric utility revenue as AI Revenue by that metric
It seems like the "Step 1, Step 2, …, Step 4" format likely predated the 1998 South Park episode, with use on Usenet for omitting dangerous or legally-questionable instructions or otherwise gatekeeping "trade secrets": https://arstechnica.com/civis/threads/where-does-the-step-st...
What’s the point of having “sovereign” weights that are worse than publicly available ones? Wouldn’t Europe be better off just keeping up-to-date on the Chinese releases? In the event of some schism requiring sovereign capability, or even if the Chinese pulled ahead and stopped releasing the weights, why would Europe be better off because of Mistral? (Or any country’s inferior sovereign effort make them better off?)
I think I understand the incentives that cause this to exist (it would be politically worse to say we’re just going to use Chinese models) but they are misguided. If sovereigns want to have valuable models, they should insist on world class, relevant ones like the Chinese have. Instead they embrace mediocrity in the name of sovereignty.
In the views of most Europeans, that schism already happened.
Europe was perfectly happy to rely on US software and services for decades. None of the large US tech companies would be nearly as profitable if they hadn’t had a whole continent of wealthy customers, and no competition.
I don’t think Americans are realizing yet how much has changed for us the past two years.
Can I still get Google Adsense payments when Google has my European bank account and address if there are sanctions or something like that? Can I still login to my Cloudflare account and manage a domain registered with them? Will I lose access to Outlook or Gmail? Should I rely on Claude at all?
But the problem isn't if I can bypass sanctions, use VPNs, etc, or even if a company wants or can legally do that... it's the fact I'm asking these questions at all. It's not something I'd ask 15 years ago about US companies.
It doesn't happen overnight, but the adversarial behaviour of the current administration and the tariffs really changed the perspective about America on many Europeans.
The discussion about building European alternatives had never been so mainstream. If and once they emerge, I think the shift will happen. But let's see
It's a slow shift, but it's coming. For instance, Airbus already picked a French AWS replacement (Scaleway). It won't be all, it won't be tomorrow, but "what happens if we get tariffed / they invade Greenland" is already part of everyone's disaster planning.
2 years is not a long time. What I’m telling you is that sentiments have changed dramatically, and every government and company board across the continent is taking actions to position itself according to those sentiments.
You won’t see the full effect of that for at least a decade, but that doesn’t mean it’s not happening.
The reality seems to be that for decades Europe gave only lip service to decoupling from American tech infrastructure, but in the last couple of years America has gone from being seen as a strong ally to being a major risk.
It will take time to move. Frankly as an American I hope it takes a long time and we get our shit together and rebuild our alliance with Europe. But it’s possible that the damage is not reversible in the next couple of decades and Europe will accelerate their decoupling. It’s also possible we continue to slide into imperialist authoritarianism (and Europe definitely accelerates their decoupling).
This is what I mean: Americans don’t understand the immense dividends they have enjoyed from being the defacto symbol of “progress” in the 20th and 21st centuries so far. American solutions were chosen by European customers because they were reliable trading partners with an air of modernity. Homegrown was seen as the antithesis to leapfrogging into the future.
People celebrated when McDonald’s came to their country or town. Not anymore.
I'm not sure if it's really a "trump agenda". The ideas didn't start with him, the people who voted for said ideas aren't going to suddenly disappear, and this is his second term.
It looks more like part of the new normal from the outside, and that's why things are changing. A country can't go from reliable to unreliable partner every 4 years without any consequences.
Even the people that voted for Trump didn't sign up for this, which is why his approval ratings are terrible even by historic standards. He didn't run on a platform of "attack Canada and suck up to Putin". Most Americans are very pro-European. Most of the voting population has living memory of the cold war.
As to what happens post-Trump, it's anyone's guess. I expect a hard reset. Markets have short memories, as they say.
Not to the same extent, but during his first term, his administration had issues with Canada, Mexico, Europe/UK, NATO, China, etc. Tariffs were a thing back then already. And he was elected again. Clearly it wasn't a huge issue for those who are bothered to go out and vote.
In any case, for everyone else, it's the second time in 10 years that they have to play the waiting game until it's over. It's good to be optimistic, but when you're in charge of a country, you can't just plan for the best and rely on hope alone. Same if you run a company or and even with services you as individual use.
Regarding memories, I don't know about the markets, but in Europe I think it will take a few years for everyone to forget that in 2026 Denmark had to transfer blood reserves to Greenland and that different European countries deployed tripwire forces because of... the US of A. Do you know who triggers similar responses? Russia, when they start saying that their borders are fluid.
2 things here, one is related to benchmaxxing, another to being good enough
1. Mistral isn't benchmaxxing. That doesn't mean they're better, but it does mean the benchmark gap not a good reflection of the actual gap
2. I think the "world class or nothing" framing mixes general capability with system capability.
Most deployments don't need AGI
In RL you need a model that's reliably good at one or two things, thats it.
Example:
Case of a hospital flooded in emails.
You make a system that decides which patient emails needs a human and drafts replies for the rest.
If a sovereign model is good enough at that, and you can run it on a hospital's own servers under EU jurisdiction, the frontier gap part has zero importance
Who cares about "beats DeepSeek / GPT11 / Claude Fairytale 8.9"
Not sure if you are European, but in EU it's a bit taboo to even talk about this in this manner. We like to spend a lot of money to make sure we finish last.
Any model, even an open weight one, is fundamentally an encoding of a way of viewing the world.
What kind of "alignment" are AI labs optimizing for? Ideological alignment is the full term, self-censored into something more technological-sounding.
Every model has people behind it rating what it should and shouldn't say. Every time you ask a model and trust its answer, you become ever-so-slightly ideologically indoctrinated.
I don't want my model to reflect the views of American oligarchs or Chinese cadres. I want European values of enlightenment and humanitarianism to be the default and that's why the sovereign part is important.
Is there a specific concern you have, and what kind of performance penalty is it worth to you on say coding tasks?
Conceptually, sure I understand, but in practice it currently seems like it amounts to just using a worse model without getting anything in return. And if some hypothetical alignment to European values is important, it seems like putting the necessary effort into building a model that’s actually competitive but has this alignment is the solution, rather than accepting an inferior one.
I think for coding tasks, it probably doesn't matter as much. I'm more worried about stuff like chatbots subtly pushing or normalizing a certain world view.
I think I'm not the only one that's had this revelation, recall the Llama4 announcement. [0]
> It’s well-known that all leading LLMs have had issues with bias—specifically, they historically have leaned left when it comes to debated political and social topics. This is due to the types of training data available on the internet.
> Our goal is to remove bias from our AI models
This goal of course is self-defeatingly impossible to achieve. There is no unbiased, there's always only an unbiased relative to the bias of the observer.
Phrased differently, models weren't right leaning enough for the American oligarchy class and they publicly shared their desire to change the ideology they perpetuate.
I think an important perspective to keep in mind here is Zizek, the philosopher who's dedicated his life's work to the functioning of ideology.
> I already am eating from the trashcan all the time. The name of this trashcan is ideology. The material force of ideology - makes me not see what I'm effectively eating. It's not only our reality which enslaves us. The tragedy of our predicament - when we are within ideology, is that - when we think that we escape it into our dreams - at that point we are within ideology.
-- Slavoj Zizek
My concern is that both the Americans and the Chinese will be aligning their models more and more ideologically and that they will function as the perfect propaganda machine - surface-level objective and unthreatening, but answering every question asked from a world view decided elsewhere.
For some tasks (like aforementioned coding), that won't matter and there we can use whatever model is most capable. But as we outsource more and more of our thinking to AI and use it more and more to educate impressionable young people, not having our own models will mean not getting a say in how societies views are shaped and perpetuated.
Imagine for example, the question: "What caused the French revolution?" There's many answers that might be technically correct. Which ones get emphasized is where ideology lives and gets perpetuated.
Since there is no unideological, the best we can do is a pluralism of ideologies building LLMs. To not build one representing your views is for your views to not be represented.
The training data and knowledge is the edge, you need to build that up and maintain it. And of course mine everything you can from the American and Chinese models, like they mined everything from the internet / films / music / games etc.
China's model for decades has been to do it cheaper and then do it better for cheaper. Just accepting this makes you an economic vassal state. We've seen this play out over decades now with other industries. I'm not blaming China for this approach, but if you want to stay relevant, then you need to compete.
You don't need to view China as some scary boogeyman who's going to use AI to attack you or w/e the current conspiracy is. They just need to continue peacefully outperforming while everyone else gets fat and lazy.
We know it's possible to put backdoors into LLMs, we don't have reliable ways to detect them without direct support from whoever inserted it.
Europe is less-worse-off with open weights than with… I guess it's weights-as-a-service? WaaS? The thing Anthropic and OpenAI do.
But that's not enough. As recently demonstrated, being just a few months behind with the power differential between defending with an open weight model while being attacked by a leading model, means losing absolutely.
I do not know if this holds going forward or not. It's not inconceivable that we're just about to get models that make unhackable code, using all the things software developers keep saying you need to do if you really care about security.
But anyone concerned about sovereignty can't bet the farm on this possibility. For the moment, it looks like it's a national security matter to ensure at least core state functionality (including core private sector logistics) gets the absolute best attention money can buy, and that the absolute best money can buy ("can buy" does a lot of heavy lifting here) is currently LLMs, and it's important those LLMs aren't going to get cut off by arbitrary whim like Mythos was, and it's important that those LLMs don't have backdoors like we can't rule out anyone else's from having.
This is true even if Europe was only defending from Russian cyberwarfare and didn't need to plan for the president of the country in which Mythos was developed, attempting to annex two NATO states.
This misunderstands what it means to be better at writing. Root cause here is that writing is how humans communicate ideas. An LLM written thing isn’t doing that, it’s something equivalent to copy/pasting a Wikipedia article. Even if the writing gets better, it will still be necessarily empty when expanding anything that isn’t completely specified.
LLMs seem already to be pretty good at translating, where you already have something fully written and are changing the language. It’s when they get rough ideas and fill in the gaps you get the empty prose they are known for.
> Root cause here is that writing is how humans communicate ideas.
Writing is a way humans communicate ideas. It's not the only way humans communicate ideas. And now, it's not only humans who communicate ideas, we just saw with the OpenAI HuggingFace hack how AI agents were able to communicate amongst themselves by using various hacked websites to opportunistically write notes for later agents to use.
All that "X is a thing only humans do" type of circular definitions will buy you, is to expand the definition of what humanity is. And I doubt that's really what you think.
I think you're missing the context, which isn't musing on what it is to be human and can machines think, but the purpose of Joe Smith's LinkedIn account expressing that he has ideas about X is to convey the message that Joe Smith is interested enough in X to venture an opinion on it, and Joe Smith has actually just set up an automated process to generate content without thinking about X, that perverts the purpose of communicating that Joe Smith is interested enough in X to propose the following ideas he been thinking about...
Similarly if Joe's contribution to his long form "idea" is a couple of bullet points, a program trained on flowery phrasing and a weighted average of everyone else's ideas isn't communicating Joe's thoughts on the topic, it's just adding words.
The debate on whether Claude actually thinks or not is orthogonal to the fact that outsourcing your "thought leadership" to it is avoiding thinking or leading. If I want to know how Wikipedia or Claude summarise wider human thought about the topic, I can find their websites thanks
I get the context, but then the comment I'd replied to would have said that humans get better at communicating their ideas by writing their ideas on their own, rather than that writing is some kind of human-only mode of exposition. It's not.
And nor is an LLM generating text just "copy/pasting a Wikipedia article", you'd think people on HN would be smarter than that at least.
If all Joe Smith is going to do is cat $(which claude) to his LinkedIn, then he'll deserve the poor results he gets from it, but we shouldn't mistakenly say that this will be because LLMs simply cannot write. It would be just as dumb for Joe Smith to do with with a professional human ghostwriter.
> An LLM written thing isn’t doing that, it’s something equivalent to copy/pasting a Wikipedia article.
Nonsense. I’m not sure whether this was ever an appropriate description of what LLMs do, but either way, they have obviously moved way, way beyond that.
Absolutely agree. A modern agent is a sophisticated tool that can link multiple sources together to create a coherent piece of information hyper-specified for an audience. However, if humans have outlines that are expandable by LLMs it makes the most sense to do that as late as possible.
I don’t like reading AI writing either but I’m sure that’s a transitory period. Single prompt text expansions are unlikely to be useful because they’re late-bindable. You could give the original to me and I might be able to understand better.
But a series of steering prompts with various sources brought in is a different story. At that point it’s just a question of whether the agent can put together good information and their current inability to do so is unlikely to mean an inherent problem.
Some kind of UI affordance for this might help: with the agent emitting tags that allow for auto-folding or expansion in a way that allows both concise text and exposition when required by the reader. Mechanical sympathy, but for code executing on a human: good old human sympathy if you will
Isn’t the current risk due to how AI is configured, like giving it a full set of tools and internet access and a goal to hack stuff?
If we think we need laws or gate keeping, why isn’t it at this level? I already can’t ddos someone or fuzz their server or whatever right, I imagine if I threw equivalent compute at old school hacking I’d just get arrested.
The quality of the “frontier” model doesn’t really matter, they just generate transcripts, they can take no action.
If this was real they’d be calling on people to stop hooking them in to “dangerous” harnesses as opposed to pausing research. But it’s not.
reply