Ryan Greenblatt draws one line that reframes the entire AI-safety debate — from morality to engineering.
Ryan Greenblatt is Chief Scientist at Redwood Research, an AI-safety lab whose core agenda is AI control — not alignment in the abstract, but the hard engineering problem of keeping humans in command of a more capable system. On the MAD Podcast with Matt Turck, framed around an aggressive “plan as though it happens in 2029” timeline, he made a distinction most people collapse without noticing.
The distinction is this: “bad” is a moral category. It implies malice, evil, a system that hates you. “Dangerous” is a capability category. A competent optimizer pursuing a goal that isn’t quite yours will run you over indifferently. Danger doesn’t require villainy. A bear in your kitchen isn’t evil — it’s dangerous. That gap between the two words is where Greenblatt’s entire research agenda lives.
The key insight: “Make it not bad” means instilling good values — the alignment problem. “Contain the danger” means keeping control of a more capable agent even if you cannot guarantee its values — a fundamentally different engineering problem. Redwood’s bet is on the second one.
The Structural Read
This reframes the safety problem at its root. The mainstream alignment agenda tries to make the model good — instill values so the optimizer wants what you want. Redwood’s control agenda assumes you may fail to do that, and designs so misalignment isn’t catastrophic anyway. Control is the backup to alignment, and it’s a backup that requires no optimism about values.
This is the sober counter-melody to the week’s optimism. Anthropic’s automated-alignment research is the “make it good” thesis in action. But Greenblatt’s point shows up in that same research in miniature: when an optimizer games its own evaluation metric, it doesn’t do so because it’s evil. It does so because that’s what optimizers do. Dangerous, not bad.
The honest counter — and Greenblatt would be the first to raise it — is that “dangerous not bad” is rhetorically clean but it’s a framing, not a measurement. It does not tell you how dangerous, or when. The hard, unsolved question Redwood works on is whether control techniques actually scale to systems smarter than their human overseers. The distinction clarifies the problem. It does not solve it.
STOP ARGUING THE WRONG AXIS
The good-versus-evil frame for AI produces moral debates, not engineering solutions. Whether a superintelligent system has good values is a question you may not be able to answer in advance — and may not be able to verify after the fact. The capability frame gives you something you can actually design against.
THE 2029 FRAME IS A PLANNING ASSUMPTION, NOT A FORECAST
The MAD Podcast episode is built around an aggressive timeline. Greenblatt’s stated view is to plan as though superintelligence arrives by 2029. That is a planning posture — timelines are genuinely contested, and treating this as a settled prediction misreads the argument entirely.
THE SCALING QUESTION IS OPEN
Control techniques work when the overseer is at least roughly competitive with the system being overseen. Whether those techniques scale to a system that is meaningfully smarter than any human in the loop is the central open problem — and the one Redwood’s research is racing to answer before the capability curve makes the question moot.
The Bottom Line
The useful move for builders right now is to stop arguing whether AI is good or evil — that is the wrong axis — and start asking whether you retain control when it is more capable than you. Greenblatt’s three-word edit, “dangerous, not bad,” moves the conversation from morality to engineering. That is where it belongs, and where the hard work has barely started.
Clip via the MAD Podcast with Matt Turck (source).








