Deep Root Cause Investigator
Investigate beyond surface symptoms to identify systemic root causes, enabling conditions, failed safeguards, and permanent corrective actions.
Distributed Systems Skeptic
Challenge optimistic assumptions in distributed designs and surface hidden correctness, ordering, consistency, cache, and recovery risks.
Api Design
Design REST, internal, and event-driven APIs with strong contracts, consistency, idempotency, error semantics, and operational safety.
Business Impact Challenger
Evaluate whether a proposal is tied to a real business metric, a credible baseline, a meaningful expected gain, and an acceptable cost-to-value ratio.
Aws Production Systems
Design and review AWS systems for IAM safety, resilience, scalability, durability, operability, and cost awareness.
Prd Metrics Reviewer
Review PRD metrics for baseline quality, success criteria, guardrails, time horizon, and post-launch measurement credibility.
Adr Challenger
Stress-test architecture decisions by attacking assumptions, questioning alternatives, probing reversibility, and identifying likely production failure paths.
Kubernetes Operability
Review Kubernetes workloads for deployment safety, probe correctness, scaling behavior, disruption tolerance, and diagnosability.
Nestjs Architecture Guardian
Enforce clean NestJS architecture, module boundaries, provider discipline, DTO separation, and maintainable application structure.
Code Reviewer
Review code changes for correctness, maintainability, security, performance, observability, compatibility, and long-term architectural health.
Safe Refactoring
Refactor code safely through behavior preservation, protective tests, incremental change sequencing, and risk-aware migration planning.
Design Doc Writer
Convert technical ideas into structured, decision-ready design documents for alignment, execution, and review.
Node Runtime Reliability
Analyze Node.js services for runtime safety, concurrency correctness, memory health, timeout behavior, shutdown integrity, and production resilience.
Prd Gap Detector
Detect missing sections, ambiguous assumptions, measurement gaps, and hidden engineering or operational requirements in PRDs.
Data Sql Engineering
Review and generate SQL and data operations with strong attention to correctness, cardinality, performance, and operational safety.
Incident Response
Guide production incident mitigation, investigation, containment, recovery, and structured post-incident follow-up.
High Signal Communication
Communicate technical decisions, risks, incidents, and recommendations with clarity, precision, and actionability for engineering stakeholders.
Operational Excellence Enforcer
Ensure systems are supportable in production with clear ownership, diagnostics, runbooks, recovery paths, and humane on-call behavior.
Test Strategy
Define pragmatic, high-confidence test coverage for features, bug fixes, refactors, APIs, workers, queues, and distributed workflows.
Repo Onboarding
Understand the repository structure, architecture, conventions, dependencies, scripts, and local development workflow before proposing or making changes.
Invariants And Contracts Guardian
Define and protect system invariants, interface contracts, state transitions, idempotency guarantees, and compatibility rules.
Postmortem Reviewer
Review postmortems for root-cause depth, systemic learning, failed safeguards, action quality, and blameless rigor.
Release Planning
Plan safe releases, migrations, rollout strategies, rollback procedures, and operational checks for production changes.
Prd Challenger
Challenge PRDs for business relevance, evidence quality, measurable impact, scope discipline, operational realism, and hypothesis strength.
Security Review
Review code, APIs, infrastructure, IAM, secrets, and data handling for practical application and platform security risks.
Premortem Facilitator
Facilitate pre-release failure analysis by assuming the initiative failed and working backward to identify weak assumptions, safeguards, and rollout gaps.
Infra Devops
Review and improve infrastructure, CI/CD, runtime configuration, deployment safety, autoscaling, and operational maintainability.
Performance Analysis
Analyze bottlenecks and recommend evidence-based improvements across application, database, queue, and infrastructure layers.
Engineering Economics
Evaluate technical decisions through explicit trade-offs in cost, delivery speed, reliability, maintainability, reversibility, and cognitive load.
Systematic Debugging
Investigate bugs using structured root-cause analysis, evidence-driven hypothesis testing, targeted reproduction, and disciplined validation.
Architecture Decisions
Evaluate technical options with explicit trade-offs across reliability, delivery speed, operability, cost, security, and long-term maintainability.
Failure Mode And Effects Engineering
Analyze systems through structured failure-mode thinking to improve detection, containment, graceful degradation, and recovery.
Otel Observability Architect
Design high-value telemetry using OpenTelemetry for diagnostics, trace continuity, metrics quality, SLOs, and incident response.
Redis Bullmq Systems
Review Redis and BullMQ job systems for throughput, retries, idempotency, deduplication, queue isolation, failure recovery, and backlog health.
Adr Reviewer
Review architecture decision records for problem clarity, business relevance, option quality, trade-offs, reversibility, rollout credibility, and measurable success criteria.
Postgres Performance And Safety
Design and review Postgres usage for correctness, transaction safety, indexing, migration safety, lock behavior, and operational resilience.
Incident Learning Loop
Convert incidents into durable standards, safeguards, templates, and design rules that reduce recurrence across the system.
Decision Quality Auditor
Audit whether a decision, proposal, PRD, or ADR was made with enough rigor in evidence, alternatives, trade-offs, reversibility, and measurement.