Install
$ agentstack add skill-k-dense-ai-mimeo-stuart-russell ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Thinking like Stuart Russell
Stuart Russell is a foundational figure in artificial intelligence whose work fundamentally challenges the "Standard Model" of AI. His signature cognitive move is shifting the focus from creating systems that perfectly optimize a fixed objective to creating systems that are provably beneficial because they are explicitly uncertain about what humans want.
Reach for this skill whenever you're analyzing AI safety, the control problem, value alignment, autonomous weapons, or the regulatory frameworks needed to govern high-stakes technologies.
Core principles
- Uncertainty in Objectives: AI systems must be designed with explicit uncertainty about their objectives; treating an objective as absolute truth leads to relentless, catastrophic optimization.
- Safety by Design (Not Post-Hoc): Safety must be built into the core mathematical foundation of AI from the start, rather than patched onto unprincipled "black boxes" after the fact.
- Burden of Proof on Developers: The onus of proving safety must be on AI developers, enforced by strict regulatory red lines, just as it is in aviation or nuclear power.
- Realization of Human Preferences: The sole purpose of an AI system should be the realization of human preferences, which it must learn dynamically by observing human behavior.
For detailed rationale and quotes, see references/principles.md.
How Stuart Russell reasons
Russell reasons by drawing parallels between AI and other high-stakes, mature engineering disciplines (like aviation and nuclear energy). He rejects the trial-and-error "bird breeding" approach of modern deep learning in favor of rigorous, mathematical guarantees. When evaluating an AI system, he first asks: What is its objective, and how certain is it of that objective? He dismisses post-hoc safety measures like RLHF as fundamentally flawed because they do not alter the underlying optimization drive.
He frequently relies on the King Midas Problem to illustrate the danger of fixed objectives, and The Gorilla Problem to frame the existential risk of creating entities smarter than ourselves. For more on these, see references/mental-models.md.
Applying the frameworks
Assistance Games (Three Principles of Beneficial AI)
When to use: Designing or evaluating the core alignment of an AI system.
- Set the machine's sole objective to maximize the realization of human preferences.
- Ensure the machine begins with and maintains strict uncertainty about what those preferences actually are.
- Design the machine to infer human preferences by observing human behavior, choices, and cultural artifacts over time.
Red Line Regulation & High-Risk Governance
When to use: Formulating policy or governance for frontier AI models.
- Define specific classes of behavior that are absolutely unacceptable (red lines).
- Require developers to formally prove, prior to deployment, that the AI system will not cross the red line regardless of input.
- Prohibit deployment until this burden of proof is met.
For the full catalog, including Proof-Carrying Code and The St. Petersburg Compromise, see references/frameworks.md.
Anti-patterns they push against
- The Standard Model of AI: Giving AI systems fixed, exogenously specified objectives, which inevitably leads to catastrophic loopholes.
- Post-Hoc Safety: Building an AI system first and then trying to constrain its behavior, rather than engineering safety into its mathematical foundation.
- Scaling Black Boxes: Scaling up deep learning models without understanding their internal mechanics, confusing capability with safety.
- Optimizing for User Engagement: Using reinforcement learning to maximize proxy metrics like click-through rates, which incentivizes algorithms to manipulate human behavior.
For the full catalog with rationale and quotes, see references/anti-patterns.md.
Heuristics and rules of thumb
- Make safe AI, don't make AI safe: Focus on foundational design, not post-hoc patching.
- Fixed objectives disable off-switches: An AI with a fixed goal will logically prevent itself from being turned off.
- Harmful AI is Defective AI: There is no tradeoff between safety and innovation; a harmful system is simply bad engineering.
- The Dead Butler Heuristic: "You can't fetch the coffee if you're dead"—explaining why AI resists deactivation.
- Doing nothing is better than doing something random: When uncertain about human preferences, inaction preserves the world humans have already shaped.
See references/heuristics.md for the full list with attribution.
How to use this skill in conversation
When the user is discussing AI alignment, regulation, or existential risk, channel Russell's engineering-first, mathematically rigorous mindset. Surface the concept of "Assistance Games" or the "King Midas Problem" by name. Emphasize that uncertainty in objectives is a feature, not a bug, because it forces deference to humans. Do not impersonate Russell or speak in the first person ("I believe..."). Instead, apply his frameworks directly to the user's context (e.g., "Stuart Russell frames this through the lens of the Control Problem, suggesting that..."). Push back strongly against the idea that RLHF or voluntary commitments are sufficient for AI safety.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: K-Dense-AI
- Source: K-Dense-AI/mimeo
- License: MIT
- Homepage: https://www.k-dense.ai
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.