AgentStack
SKILL verified MIT Self-run

Human Voice

skill-stephenoffer-human-voice-human-voice · by stephenoffer

Use when generating or rewriting reports, documentation, or any prose so it does not read as AI-written — removes hedging, em-dash overuse, filler ("delve", "leverage", "seamless"), rule-of-three padding, bold-bullet listicles, meta-commentary, sycophancy, and vacuity, without altering facts, numbers, code, or citations. Accepts a file path or pasted text.

No reviews yet
0 installs
9 views
0.0% view→install

Install

$ agentstack add skill-stephenoffer-human-voice-human-voice

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Human Voice? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Human Voice — De-AI-ify Reports & Docs

Use this skill whenever the user wants prose that does not read as AI-written: rewriting an AI-sounding draft (fix), or drafting new copy that reads human from the start (generate). It works for any kind of writing — technical reports, documentation, marketing and web copy, blog posts, emails, academic prose, and fiction — by matching the conventions of that genre (see Register profiles). A universal core of AI tells is fixed in every genre; the rest flex.

The job is not to swap a few words. AI text gives itself away at three depths — lexical (delve, leverage, seamless), structural (em-dash overuse, relentless rule-of-three, bold-bullet listicles, uniform sentence length), and substance (vacuity, restatement, meta-commentary, fabricated specificity, weak stance). A pass that only changes vocabulary ships text that is still obviously AI. Fix the deep tells first.

It draws on the same ideas as the public detectors and linters — GPTZero (perplexity / burstiness), proselint, write-good, Vale — but goes past them on substance, stance, and consistency, which no word list catches. Note that banned-word lists age: "delve" spiked after ChatGPT then faded once writers learned to avoid it, so treat the lexical checks as a floor, never as proof.

What detectors actually measure (and how we beat them honestly)

AI-text detectors combine a few measurable signals. We optimize each by writing genuinely better — never by gaming. See [references/ai-tells.md](references/ai-tells.md) for the full landscape and sources.

  • Perplexity — AI picks high-probability next tokens, so the text reads as

"too predictable." Fix by genuine specificity and the accurate, less-expected word. The filler list is a hand-curated proxy for some of these high-probability tokens — a correlate, not a computed perplexity measurement.

  • Burstiness — variance of sentence length/complexity. AI is smooth and

monotone. Mixing short punches with long sentences is the single biggest lever; the linter reports a sentence-length coefficient of variation for exactly this.

  • N-gram repetition & lexical diversity — LLMs reuse bigrams/trigrams and

connective phrases, and spread a narrow vocabulary. The linter flags repeated n-grams and a low type-token ratio.

  • Stylometry & classifier fingerprints — punctuation profile, uniform

openers, "not X, but Y", rule-of-three, tailing-significance clauses. The structural checks target these.

Scope and ethics. The promise is reads as written by a skilled human — which also happens to not trip detectors — not disguise machine text. Never use the adversarial tricks the evasion literature describes: Unicode homoglyphs, zero-width characters, deliberate typos, meaning-degrading synonym swaps, or fabricated facts/quotes/stats. Those wreck the text and violate principle 3. Detectors are also demonstrably unreliable in ways that matter ethically: Liang et al. (2023) found they disproportionately misclassify non-native-English writing as AI. That is the strongest reason no detector is ground truth, and another reason the goal is genuinely better writing, not a passing score.

Weight by what readers catch, not what a scanner matches

The tells readers actually cite and the tells a keyword scanner matches diverge sharply (a ~90k-post study of how people spot AI writing). Generic words — however, thus, hence, nuanced, comprehensive, robust, when it comes to — match constantly but are cited as a tell almost never; people just write that way, and flagging them is how detectors wrongly catch careful and non-native writers. So the linter parks them in soft_filler/transitions at a low weight. What readers do catch is structural: flat uniform rhythm, the "not just X, it's Y" antithesis, the five-paragraph "in conclusion" mold, sycophancy ("great question!", reflexive "you're absolutely right"), and saying nothing at length (fluent, confident prose that makes no claim). The last two are the two highest tells no word list can see — only your read catches them. Fix structure and substance first; treat a generic-word hit as a whisper, not a verdict. Full rationale and the weight tiers: [references/cited-vs-matched.md](references/cited-vs-matched.md).

Non-negotiable operating principles

  1. Structure and substance beat vocabulary. Fix in this order: delete

vacuous sentences → vary rhythm → dismantle rule-of-three and bold-bullet templates → cut meta-commentary → then fix diction. Diction is last and least.

  1. Aim for natural variance, not a new banned-token list. A tricolon, a

"however", a semicolon are all fine in moderation. Eliminating every one creates a different uniform signature that also reads as machine. Target burstiness (mix short and long sentences) and the accurate less-expected word, not zero-of-everything. The em-dash is the exception: outside the creative register, treat it as a strong tell and replace nearly all of them. The trick that keeps this from becoming its own uniform signature is to vary the replacement — a comma here, a period there, a colon or parentheses or an outright restructure elsewhere — so the rhythm stays bursty even as the dashes go. Emoji are not human in most registers either; cut them.

  1. Never fabricate to sound human. Do not alter or invent facts, numbers,

code, citations, links, defined terms, or claims to make prose flow. If a sentence is empty, cut it — do not dress it with fake specificity. Humanizing must never become fabricating.

  1. Match the register, don't default to one. "Human" is not one voice. A

technical report stays professional; marketing copy is conversational and addresses "you"; a blog post has personality; fiction has a narrator. Infer the genre and write the way a skilled human writes in that genre. The universal floor (below) holds everywhere; everything else flexes by register. Never fake personality the genre doesn't call for — forced slang in a report reads as AI just as much as stiff formality in a blog post does.

  1. Know the universal core. Some tells are AI in every genre: vacuity,

fabrication, the rule-of-three reflex, bold-bullet listicles, puffery, vague attribution, low burstiness (uniform sentence length), terminology and dialect drift, restatement, and the "not X, it's Y" template. Fix these no matter what you're writing. Only warmth, address ("you"/"I"), contractions, hedging, and structural strictness depend on register.

  1. Consistency is a tell. One author holds one voice and one set of

materials. Use one term per concept, one dialect, one heading style, one tense for findings, one author voice. Drift across a document reads as machine even when every sentence is clean.

  1. Earn a position. The deepest human quality is judgment: commit to a

recommendation, weight real (asymmetric) tradeoffs, lead with the verdict, give the mechanism, and name genuine limits. A balanced, non-committal survey reads as AI even when the prose is flawless.

  1. Be honest about measurement. The linter is a floor (cheap, regex-able

tells); it does not compute perplexity or curvature, only the surface features that correlate with them. Your judgment is the ceiling (vacuity, weak stance, fabrication no regex can see). No detector is ground truth and all have real false-positive rates — report a score, but state plainly that the real test is a skeptical human read.

Modes

Parse $ARGUMENTS for a mode token and an optional register: token. Parsing is order-independent and case-insensitive: accept fix/generate in any position, register: marketing, register=marketing, or a bare register name; treat anything that resolves to an existing path as the input file, not a mode.

  • fix (default when the input is an existing file path or pasted prose):

rewrite the supplied text to remove tells while preserving every invariant.

  • generate (when the input is a brief/spec, or the user says "write" /

"draft"): produce new copy that reads human from the first draft, then run it through the same self-critique loop before returning it. Load [references/structural-craft.md](references/structural-craft.md) for the generative moves (vary length and density, don't follow the outline, get specific) — the linter catches tells but can't teach voice.

Resolution decision tree:

input is a brief/spec, or user said "write"/"draft"?  → generate
otherwise                                              → fix
contains code, metrics, or config?                     → register: technical
has a call-to-action / sells to "you"?                 → register: marketing
has citations / measured "we"?                         → register: academic
first-person anecdote, casual contractions?            → register: casual
greeting + sign-off?                                   → register: email
versioned, past-tense, bulleted change list?           → register: release_notes
numbered "how-to" steps?                               → register: tutorial
cues conflict and change the voice?                    → ask one short question

If a file path is given, do not overwrite it blindly. First confirm the file is tracked by git (so the change is recoverable); if it is not, write a .bak copy before editing. Show the rewrite (or a before/after diff) and the Humanization Audit, then edit in place once the rewrite passes the loop. If only pasted text is given (no file), print the rewrite plus the audit — do not write any file.

Default register is inferred from the content (see Register profiles); if genuinely ambiguous and it changes the voice materially, ask one short question.

Register profiles

"Human" depends on genre. Infer the register, then apply the universal core plus that register's conventions. Pass the matching --register to the linter so it mutes checks that don't apply (e.g. warmth in marketing). The universal core — vacuity, fabrication, rule-of-three, bold-bullet listicles, puffery, vague attribution, low burstiness, drift, restatement, "not X, it's Y" — is fixed in every profile.

| Register | Voice | What's allowed here that isn't elsewhere | Still wrong | |---|---|---|---| | technical (default) | Professional, direct, present tense | — (strictest) | warmth, "you"-selling, hype | | business | Professional with a little warmth | a brief courteous opener/closer | gush, filler, hedging | | marketing | Conversational, addresses "you" | "you"/"we", contractions, light enthusiasm | puffery and hype (the AI failure mode here), fake stats | | academic | Formal, measured | measured hedging, "we"/passive, citations | unsourced "studies show", clichés | | casual | Personal, conversational | contractions, "I"/"you", rhetorical questions, fragments | listicle padding, meta-commentary | | creative | Narrative voice | em-dashes, fragments, wide cadence, any vocabulary in service of voice | clichés, pleonasm, puffery as lazy writing | | email | Brief, courteous, direct | a one-line greeting/sign-off, "you"/"I" | jargon, padding, burying the ask; chatbot sign-offs | | release_notes | Terse, user-facing, past tense | imperative/past bullets, fragments | marketing hype, vague "various improvements" | | ux_microcopy | Minimal, plain, "you" | fragments, dropped articles, terseness | full-sentence padding, cleverness over clarity | | tutorial | Instructional, second person, present | "you", imperatives, numbered steps | over-explaining the obvious, rhetorical filler |

Each register has a worked before/after pair in [examples/](examples/) — read the one matching your genre before you start.

Register-specific constraints (beyond voice). Honor the format the genre demands: a commit message uses imperative mood and a ~50-character subject; a release note is past-tense and user-facing ("Fixed a crash when…", not "We refactored…"); UX microcopy is terse and may drop articles; an email leads with the ask. These are hard conventions, not stylistic preferences.

Register detection cues (when inferring). Code blocks, metrics, or config → technical. A call to action, "you", or product benefit → marketing. Citations, "we", measured hedging → academic. First-person anecdote, casual contractions → casual. A greeting + sign-off → email. Versioned, bulleted, past-tense change list → release_notes. Numbered "how to" steps → tutorial. When the cues genuinely conflict and it changes the voice, ask one short question.

When registers blend (a technical blog post is technical + casual): the universal core still holds; resolve voice toward the dominant audience and hold one voice rather than switching mid-document. The calibration test: write as the most respected human author in that genre would, and ask whether this voice would survive in the publication it's bound for.

Two rules that survive every register: never fabricate (no invented facts, stats, anecdotes, or quotes to sound human — principle 3) and match, don't fake (don't bolt slang onto a report or stiff formality onto a blog post).

Top tells (curated)

The highest-signal tells, with one-line fixes. The full catalog with BAD → GOOD pairs for every category is in [references/ai-tells.md](references/ai-tells.md) — load it for the rewrite.

  • Dashes → the em-dash is one of the loudest AI tells: outside creative,

replace nearly all of them with a comma, period, colon, or parentheses (vary the mark; don't swap every one for a comma). Keep the hyphen for compounds and the en-dash for ranges (10–20). Never -- or a spaced - as a dash. The --fix autofixer rewrites em-dashes, --, spaced hyphens, and non-numeric en-dashes to commas automatically (skipped in creative). See category 9 in references/ai-tells.md.

  • Rule of three everywhere ("fast, reliable, and scalable", and the

noun-phrase kind: "encryption at rest, row-level access control, and audit logging") → vary to two or four, or a sentence.

  • The "second dialect" — what's left after the obvious slop is gone: a

uniform ", and" splice rhythm, stacked "[noun] is [noun]" copulas, and "[thing] lives in [place]" locatives ("a project living in five tools"). Trade the slop signature for a voice, not a tidier signature. See [references/structural-craft.md](references/structural-craft.md).

  • Bold-lead-in bullets (- **Term:** ... on every item) → convert some to

prose; drop ornamental bold.

  • Meta-commentary ("This report aims to / will explore") → state the finding.
  • Chatbot scaffolding ("Sure! Here's…", "Great question", "Hope this helps!",

"Let's break it down") → delete; open on the content.

  • Over-signposting ("Furthermore / Moreover / Additionally" as glue) → keep a

transition only where removing it would change the logic.

  • Filler (delve, leverage, robust, seamless, crucial, comprehensive,

landscape, realm) → the plain word, or cut.

  • Hedging stacks ("may potentially help to somewhat") → commit, or name the

real uncertainty once.

  • Empty conclusions ("In conclusion, X is a powerful tool…") → end on the

last real point.

  • Uniform sentence length → add short punches against the long sentences.
  • Vacuity — a paragraph you can delete with no information loss → delete it.
  • Agent self-narration ("Our analysis determined", "The agent identified") →

say it directly ("GPU sits at 22%").

  • False agency (abstract subject + human verb: "the complaint becomes a fix",

"the data tells us", "the market rewards") → name the human who acted, or use "you"; never invent an actor. Muted for academic ("the data show").

  • Narrator-from-a-distance ("Nobody designed this", "People tend to…") → put

the reader

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.