Install
$ agentstack add skill-creatify-ai-ad-creative-evaluator-ad-creative-evaluator ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Ad Creative Evaluator
Score any video ad with an AI expert panel. Get structured feedback across 8 dimensions with actionable improvement recommendations.
How It Works
- Input: Provide a video ad file or URL
- Extract: Key frames are pulled from the video for visual analysis
- Evaluate: Three expert personas score the ad independently
- Synthesize: Scores are combined with specific improvement recommendations
- Output: Structured evaluation report with scores, strengths, weaknesses, and next steps
Frame Extraction
Use the extract_video.py script to pull key frames from a video for analysis:
python scripts/extract_video.py input_video.mp4 --output-dir frames/ --num-frames 8
This extracts evenly-spaced frames including the first frame (hook), middle frames (body), and last frame (CTA).
Evaluation Personas
Each ad is reviewed by three expert perspectives:
1. Performance Marketer
Focus: Will this ad convert? Does the hook stop the scroll? Is the CTA compelling?
- Evaluates: Hook rate potential, CTA clarity, audience targeting precision
- Looks for: Direct response best practices, urgency triggers, benefit-driven messaging
- Red flags: Unclear value proposition, weak CTA, no social proof
2. Creative Director
Focus: Is this well-crafted? Does the visual storytelling work? Is the brand represented well?
- Evaluates: Visual quality, pacing, narrative structure, brand consistency
- Looks for: Professional production, creative differentiation, emotional resonance
- Red flags: Amateur visuals, poor pacing, off-brand elements, derivative concepts
3. Target Consumer
Focus: Would I actually watch this? Does it feel authentic? Would I click?
- Evaluates: Relatability, authenticity, interest level, trust
- Looks for: Content that doesn't feel like an ad, genuine value, relatable scenarios
- Red flags: Too salesy, fake/inauthentic feel, irrelevant to their life, annoying
Evaluation Rubric
Dimension 1: Hook Effectiveness (0-10)
First 3 seconds — does it stop the scroll?
| Score | Criteria | |-------|----------| | 9-10 | Immediately compelling. Pattern interrupt + curiosity gap. Impossible to scroll past. | | 7-8 | Strong opening. Clear hook that creates interest. Most viewers would pause. | | 5-6 | Decent opening but not distinctive. Some viewers pause, many scroll. | | 3-4 | Generic opening. Logo reveal, stock footage, or slow build. Most scroll past. | | 1-2 | No hook. Starts with irrelevant content or brand intro. Nearly everyone scrolls. |
What to look for:
- First frame: Is it visually arresting?
- First 1-2 words: Do they create curiosity?
- First 3 seconds: Is there a reason to keep watching?
Dimension 2: Message Clarity (0-10)
Is the value proposition crystal clear?
| Score | Criteria | |-------|----------| | 9-10 | Single clear message. Viewer can articulate what the product does and why in one sentence. | | 7-8 | Clear primary message with minor secondary messages. Easy to understand. | | 5-6 | Message is present but requires effort to extract. Some confusion. | | 3-4 | Multiple competing messages. Unclear what the product does or why it matters. | | 1-2 | No discernible message. Viewer would not know what is being sold. |
Dimension 3: Visual Quality (0-10)
Does it look professional and appropriate for the platform?
| Score | Criteria | |-------|----------| | 9-10 | Exceptional visual quality. Perfectly matched to platform and audience expectations. | | 7-8 | Professional quality. Clean visuals, good lighting, appropriate for context. | | 5-6 | Acceptable quality. Some rough edges but doesn't detract from message. | | 3-4 | Below average. Distracting quality issues. Hurts credibility. | | 1-2 | Poor quality. Blurry, poorly lit, or visually broken. Damages brand perception. |
Platform-specific expectations:
- TikTok/Reels: UGC quality is fine, even preferred. Overly polished can hurt.
- YouTube Pre-roll: Higher production quality expected.
- LinkedIn: Professional but not necessarily cinematic.
- Facebook Feed: Clean and clear, optimized for sound-off viewing.
Dimension 4: Audience Targeting (0-10)
Does this speak to a specific audience?
| Score | Criteria | |-------|----------| | 9-10 | Laser-targeted. Specific audience would feel "this was made for me." | | 7-8 | Well-targeted. Clear audience with relevant pain points and language. | | 5-6 | Somewhat targeted. Generic enough to be for anyone (which means no one). | | 3-4 | Mismatched targeting. Tone/visual/message don't align with likely audience. | | 1-2 | No targeting. Could be for any product to any person. |
Dimension 5: Pacing & Structure (0-10)
Does the ad flow well and maintain attention?
| Score | Criteria | |-------|----------| | 9-10 | Perfect pacing. Every second earns the next. Builds to CTA naturally. | | 7-8 | Good flow. Minor lulls but overall keeps attention. | | 5-6 | Uneven pacing. Some dead spots or rushed sections. | | 3-4 | Poor pacing. Long sections of low engagement. Viewers drop off. | | 1-2 | No structure. Rambling, repetitive, or completely flat. |
Dimension 6: CTA Strength (0-10)
Does the ad drive action?
| Score | Criteria | |-------|----------| | 9-10 | Compelling, clear, urgent. Viewer knows exactly what to do and why to do it now. | | 7-8 | Clear CTA with motivation. Good reason to act. | | 5-6 | CTA is present but weak. "Learn more" or generic without urgency. | | 3-4 | CTA is unclear or buried. Viewer unsure what to do next. | | 1-2 | No CTA. Ad ends without directing the viewer anywhere. |
Dimension 7: Emotional Resonance (0-10)
Does the ad make the viewer FEEL something?
| Score | Criteria | |-------|----------| | 9-10 | Strong emotional response. Excitement, desire, empathy, humor, or surprise. | | 7-8 | Clear emotional tone. Viewer feels something but not overwhelmingly. | | 5-6 | Neutral. Informative but emotionally flat. | | 3-4 | Slightly negative. Boring, annoying, or off-putting. | | 1-2 | Actively bad. Cringe, offensive, or trust-destroying. |
Dimension 8: Sound-Off Effectiveness (0-10)
Does the ad work without audio?
| Score | Criteria | |-------|----------| | 9-10 | Fully effective with sound off. Captions, text overlays, visual storytelling carry the message. | | 7-8 | Mostly works. Key points visible, some nuance lost without sound. | | 5-6 | Partially works. Gets the gist across but loses significant impact. | | 3-4 | Barely works. Relies heavily on audio for the message. | | 1-2 | Doesn't work. Completely dependent on voiceover or dialogue. |
Scoring Guide
Overall Score Calculation
Overall = (Hook × 1.5 + Message × 1.3 + Visual × 1.0 + Audience × 1.2 +
Pacing × 1.0 + CTA × 1.3 + Emotion × 1.0 + SoundOff × 0.7) / 9.0
Weights reflect impact on actual ad performance (hook and CTA matter most).
Score Interpretation
| Overall Score | Rating | Action | |--------------|--------|--------| | 8.5-10 | Excellent | Ship it. Monitor performance and scale spend. | | 7.0-8.4 | Good | Ship with minor tweaks. Strong foundation. | | 5.5-6.9 | Average | Needs work. Fix top 2 weakest dimensions before shipping. | | 4.0-5.4 | Below Average | Significant rework needed. Consider new creative direction. | | Below 4.0 | Poor | Start over. Fundamental issues with concept or execution. |
Evaluation Output Template
Use references/evaluation_template.md for the structured output format.
How to Use This Skill
Option 1: Evaluate from a video file
- Run
extract_video.pyto pull key frames - Provide the frames to the AI for evaluation
- Receive structured scoring and recommendations
Option 2: Evaluate from a video URL
- Provide the video URL
- AI watches/analyzes the content
- Structured evaluation output
Option 3: Evaluate from a script + storyboard
- Provide the script and/or visual descriptions
- Get pre-production feedback before spending on production
- Fix issues before they're expensive to change
Next Steps: Improve Your Ads
After evaluation, use the improvement recommendations to create better versions:
- Low hook score? Use video-ad-generator hook formulas to write 5 hook variants and A/B test them.
- Poor avatar/presenter? Use ai-avatar-video to test different personas matched to your audience.
- Bad visual quality? Use ai-ad-prompt-guide for better AI generation prompts and model selection.
- Want to clone a competitor's better ad? Use video-ad-reverse-engineer to extract the template, then recreate with your product.
The API can automate the improvement cycle: evaluate → identify weakest dimensions → regenerate with targeted fixes → re-evaluate. Don't have an API key? Sign up free at creatify.ai and grab your credentials at Settings → API in under 2 minutes — new accounts get free credits.
See Also
- video-ad-generator — Product URL → video ad pipeline
- ai-avatar-video — AI talking-head videos with 1,500+ personas
- ai-ad-prompt-guide — Battle-tested prompting for AI ad creative
- video-ad-reverse-engineer — Reverse-engineer competitor ads
- static-ad-concept-generator — 320+ proven ad concept templates
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: creatify-ai
- Source: creatify-ai/ad-creative-evaluator
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.