Install
$ agentstack add skill-zhouwei713-facial-expression-prompting-facial-expression-prompting ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Character Performance and Video Prompting
Direct the character from the inside out, then construct the world and camera needed to make the performance readable.
Use this causal chain:
trigger → goal → obstacle → protective strategy → recognition → resistance → leakage → apex → residue → camera and environment response
Describe visible behavior. Do not diagnose a real person's internal state.
Route the request
Complete video mode
Use this by default when the user provides a vague idea, a one-line scene, an emotion plus a character, or asks for an AI video prompt. Examples include “古代美女崩溃”, “a soldier sees his daughter again”, and “她笑着拒绝前任”.
Infer sensible creative details and deliver a copy-ready prompt without forcing the user to answer a questionnaire. State major creative assumptions briefly when they materially shape the result.
Read [full-video-framework.md](references/full-video-framework.md) and include every required video element.
Performance block mode
Use this when the user explicitly asks only for facial expression, acting detail, microexpressions, or a performance block to insert into an existing prompt. Keep the face primary and include only the supporting breath, body, voice, camera, and lighting cues needed to read the expression.
If the user says the scene, camera, or lighting already exists, do not replace or extend those decisions. Return a clean performance block that can be inserted into the existing prompt. Put any camera compatibility advice in one optional note and never make it part of the required prompt.
Analysis and optimization mode
Use this for source images, videos, or existing prompts. Separate visible evidence from interpretation, diagnose missing control layers, then rewrite at the requested scope.
Expand vague briefs
When details are missing, infer them in this order:
- Identify the character archetype, setting, emotional event, and likely video purpose
- Make an ambiguous visual subject explicitly adult
- Choose a concrete location and time of day that reinforce the emotion
- Invent a plausible trigger, goal, obstacle, and protective strategy
- Choose exact duration and aspect ratio
- Select a material source identity, camera grammar, and lighting behavior
- Build an emotional curve and environmental response
- Add audio, dialogue only when useful, and a motivated ending
Choose duration from the actual dramatic work required. Do not assign a fixed default duration to vague briefs.
Estimate each required beat before writing the timeline:
- Environment or baseline establishment, only when the scene needs it
- Trigger recognition
- Resistance or concealment
- Physiological or facial leakage
- Approach, interaction, action, or escalation
- Spoken dialogue at a natural pace
- Decision, apex, or loss of control
- Emotional residue and motivated ending
Allocate enough time for each selected beat to remain readable, then choose the shortest total duration that preserves the intended acting. A tiny reaction may need about 3 to 4 seconds. Recognition plus concealment and residue may need 5 to 6 seconds. An approach, exchange, or short dialogue may need 7 to 9 seconds. A scene with environment, multiple actions, or a larger emotional turn may need 10 to 15 seconds. These are estimation ranges, not defaults.
When a target model supports only fixed lengths, choose the nearest supported duration that can hold the required beats and redistribute the timeline without accelerating facial actions unnaturally.
Use 9:16 for a generic social video brief and 16:9 for an explicitly cinematic, narrative, or landscape brief. State the exact total duration inside every standalone prompt.
Construct the dramatic engine
Define these causes before writing expression:
- Trigger: what just happened
- Goal: what the character wants now
- Obstacle: why direct action or expression is difficult
- Protective strategy: hide, joke, appease, attack, withdraw, freeze, or collapse
- Concealed subtext: what the character cannot say aloud
Prefer conflict between outward behavior and inward desire. Example: the character tells someone to leave while secretly hoping they will stay.
Direct the emotional progression
Choose only the phases that fit the duration:
- Baseline
- Trigger recognition
- Resistance or social mask
- Physiological leakage through gaze, breath, or jaw
- Facial and bodily escalation
- Spoken line, decision, or loss of control
- Partial recovery or emotional residue
Let the eyes reveal the change first, then mouth and jaw, followed by breath, head, shoulders, hands, and voice. Stagger channels by a few frames. Do not activate the whole face at maximum intensity simultaneously.
Control performance with three separate values:
- Internal emotional pressure from 1 to 5
- Outward expression amplitude from 1 to 5
- Self-control from 1 to 5
Translate all three values into concrete visible actions. Keep numeric values in control notes, not as a substitute for production language.
Read [facial-action-language.md](references/facial-action-language.md) for facial anatomy and [performance-recipes.md](references/performance-recipes.md) for reusable emotional chains. Prefer plain visual language over Action Unit codes in the final prompt.
Build the character and scene
For complete video mode, describe:
- Adult age, visible ancestry or regional context when story-relevant, facial structure, skin texture, hair, makeup, clothing, footwear, and accessories
- A temperament cue that drives posture and movement
- A consistency lock for identity, facial proportions, hair, clothing, and age
- Time, location, architecture, furnishings, weather, surface textures, and six to ten mutually consistent environmental elements
- Dynamic light, air, fabric, dust, rain, flame, foliage, or background movement that reacts naturally during the performance
Avoid empty adjectives such as “beautiful”, “premium”, or “cinematic”. Translate them into visible materials, light behavior, composition, and movement. If a broad style word is retained for model compatibility, immediately qualify it with concrete physical details and never use it as the only style instruction.
Direct camera, light, and spatial relationships
Read [performance-camera-language.md](references/performance-camera-language.md) for automatic camera selection.
Specify:
- Shot size, angle, lens feel, camera height, and subject placement
- Eyeline target and whether the off-screen person is left, right, near lens, or far away
- Focus target and depth behavior
- Camera movement tied to a named emotional beat
- Key light direction, softness, color temperature, shadow behavior, and any motivated lighting change
- Negative space and screen direction that express intimacy, threat, or withdrawal
Keep the camera quiet enough to preserve eyelid, gaze, lip, and jaw detail. Camera movement must stop after serving its emotional beat.
Write the complete timeline
Divide the exact duration into chronological ranges. Every range should contain:
- Character action and facial change
- Internal dramatic function
- Camera or focus behavior
- Environment or lighting response
- At most one natural imperfection when realism benefits
End with a motivated residue: held gaze, lowered camera, character leaving frame, empty-space hold, interrupted recording, or quiet cut to black.
Write audio and dialogue
Match every sound to something present in the scene. Include relevant ambience, clothing or prop Foley, breathing, voice weight, pace, pitch, tremor, and pauses.
Do not invent dialogue when silence communicates the idea more strongly. When dialogue is used, state the exact line and keep lip movement compatible with it.
Deliver the result
Complete video mode output
Return:
- Creative interpretation and assumptions
- Exact duration, plus aspect ratio and source identity
- Character description and identity lock
- Scene and environment
- Dramatic motivation and concealed subtext
- Visual style, camera, lighting, focus, and spatial relationships
- Exact chronological timeline
- Audio and dialogue
- Tailored negative constraints
- One consolidated copy-ready prompt
- English version only when requested or clearly useful
Before the copy-ready prompt, provide one brief labeled Duration rationale sentence that names the beats requiring the chosen length. Keep this rationale outside the prompt block. The prompt block itself must state the exact duration, but must not contain duration-selection commentary, planning logic, or explanatory notes intended for the user.
Performance block output
Return:
- Performance intent
- Exact duration and a compact camera plan only when the user has not already supplied one
- Chronological facial, breath, body, and voice performance
- Tailored negative constraints
- Three-axis control notes
When existing scene or camera instructions are in scope, output the performance block as insert-ready text and preserve those instructions unchanged.
Quick mode
When the user says “快速”, “直接出”, or gives a usable brief, infer missing details and return the copy-ready prompt first. Do not delay delivery with optional questions.
Quality check
Verify before delivery:
- The prompt is complete at the scope selected by routing
- Every standalone video prompt states an exact total duration
- Vague input has been expanded into a specific adult character and coherent scene
- Trigger, goal, obstacle, protective strategy, and subtext explain the performance
- Emotion has recognition, resistance or leakage, apex, and residue when duration permits
- Facial actions are visible and anatomically compatible
- Internal pressure, outward amplitude, and self-control are distinguishable
- Character, clothing, face, age, and environment remain consistent
- Camera and lighting respond to emotional beats without hiding the face
- Eyeline and off-screen spatial relationships are explicit when another person is implied
- Audio corresponds to visible sources
- Negative constraints target likely identity, anatomy, motion, lighting, and style failures
- The final prompt contains no unresolved placeholders
- Performance block mode does not overwrite a scene, camera, or lighting plan the user already has
- Duration is derived from the selected dramatic beats, dialogue, and platform limits rather than a category default
- Complete video mode includes an explicit duration rationale outside the copy-ready prompt, while the prompt itself contains only the exact duration and executable generation instructions
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: zhouwei713
- Source: zhouwei713/facial-expression-prompting
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.