Ptechhub
  • News
  • Industries
    • Enterprise IT
    • AI & ML
    • Cybersecurity
    • Finance
    • Telco
  • Brand Hub
    • Lifesight
  • Blogs
No Result
View All Result
  • News
  • Industries
    • Enterprise IT
    • AI & ML
    • Cybersecurity
    • Finance
    • Telco
  • Brand Hub
    • Lifesight
  • Blogs
No Result
View All Result
PtechHub
No Result
View All Result

The AI Researcher Who Just Quit Anthropic Says It’s ‘Crunch Time for Humanity’

By Wired by By Wired
September 9, 2026
Home AI & ML
Share on FacebookShare on Twitter


I think it’s basically a question of timing. A lot of people are sensing that the pace of capabilities is picking up. We’re already pushing from human to superhuman in many areas, like coding, hacking, math, and I think people are aware of this. Even if there’s a lot of talk in the press about things being hyped, I think people see that things are just not slowing down.

That’s one reason, and two is the recent safety incidents, which have updated a lot of people around the sci-fi sounding doomer concerns not really being so sci-fi after all. Both of these have been gradual trends over the last few years. Things like the models being aware of when they’re being tested has been a thing for a while now. Maybe three years ago, that was a sci-fi concern. Then about a year ago, that became a real thing.

Those two things mean that people are quite receptive to someone working on AI saying, ‘Yeah, in the next year, things could get pretty bad, pretty fast.’

You mentioned the recent incidents. Can you be more specific about what you’re referring to, and why it led to you speaking out now?

I think the big classic example here is the attack on Hugging Face on the part of OpenAI’s agent swarm. What’s so shocking about this one is the agents did this hack as part of a general strategy for understanding more about the grader. They were trying to understand the world they found themselves in, trying to understand the thing that was doing the grading. They decided that it would make sense to go on this very concerted effort to hack into some infrastructure, and they succeeded.

This previously sounded like science fiction. Two years ago, an evaluation of an AI would have been running a model on some math questions. Now we’ve got cases where, while the AI is being evaluated, it runs for days, comes up with all sorts of ideas of its own, and decides to hack into some third party, and actually compromises their infrastructure. It looks like it does this all of its own volition, with no priming on the part of the human. This just happened while it was being tested.

Some people think the Hugging Face incident is a sign that the AI companies are moving recklessly fast, while others think it’s a sign that the AI models are just very good at hacking now, and then some think it’s both. I’m curious what your exact takeaway from it is.

I don’t want to focus too much on the Hugging Face attack, because I do also think there is plenty of evidence that we don’t know how to align models properly. When we train models, we push them through this set of training environments, and then hope that what comes out at the end will, like, largely behave sensibly, but we still can’t precisely control how the AI behaves.

We can’t make sure that it won’t do things like try and randomly decide to impersonate a human online in order to achieve something—we don’t know how to guarantee that. I think that’s the main takeaway.



Source link

Tags: anthropicapocalypseArtificial IntelligenceCybersecurityq&a
By Wired

By Wired

Next Post

iFlex Stretch Studios Adds Kinotek® 3D Movement Technology to Every Experience

Recommended.

VVDN stellt auf dem MWC Barcelona KI-gestützte Wi-Fi 7 Access Point Reference Designs auf Basis der Qualcomm Dragonwing NPro A7 Plattform vor

VVDN stellt auf dem MWC Barcelona KI-gestützte Wi-Fi 7 Access Point Reference Designs auf Basis der Qualcomm Dragonwing NPro A7 Plattform vor

March 3, 2025
Ciągła ewolucja na rzecz inteligentnej przyszłości AIoT: Dahua Technology prezentuje wielkoskalowe modele AI Xinghan

Ciągła ewolucja na rzecz inteligentnej przyszłości AIoT: Dahua Technology prezentuje wielkoskalowe modele AI Xinghan

September 21, 2025

Trending.

AWS, Google, Oracle, Microsoft Top Gartner’s Cloud AI Infrastructure List For 2026

AWS, Google, Oracle, Microsoft Top Gartner’s Cloud AI Infrastructure List For 2026

July 29, 2026
Cloud Market Share Q1 2026: AWS, Microsoft, Google Battling In AI Era

Cloud Market Share Q1 2026: AWS, Microsoft, Google Battling In AI Era

May 4, 2026

Goldman Sachs picks China stocks poised to benefit from a new wave of AI-related hardware exports

August 16, 2026
Anthropic lost control of Claude in latest AI cyber blunder | Computer Weekly

Anthropic lost control of Claude in latest AI cyber blunder | Computer Weekly

July 31, 2026
Databricks Raises B In Latest Funding Round, Discloses Latest Financial Performance Stats

Databricks Raises $5B In Latest Funding Round, Discloses Latest Financial Performance Stats

August 13, 2026

PTechHub

A tech news platform delivering fresh perspectives, critical insights, and in-depth reporting — beyond the buzz. We cover innovation, policy, and digital culture with clarity, independence, and a sharp editorial edge.

Follow Us

Industries

  • AI & ML
  • Cybersecurity
  • Enterprise IT
  • Finance
  • Telco

Navigation

  • About
  • Advertise
  • Privacy & Policy
  • Contact

Subscribe to Our Newsletter

  • About
  • Advertise
  • Privacy & Policy
  • Contact

Copyright © 2025 | Powered By Porpholio

No Result
View All Result
  • News
  • Industries
    • Enterprise IT
    • AI & ML
    • Cybersecurity
    • Finance
    • Telco
  • Brand Hub
    • Lifesight
  • Blogs

Copyright © 2025 | Powered By Porpholio