AgentStack
SKILL verified MIT Self-run

Paper Illustration

skill-raja21068-autoresearch-paper-illustration · by raja21068

Generate publication-quality AI illustrations for academic papers using Gemini image generation. Creates architecture diagrams, method illustrations with Claude-supervised iterative refinement loop. Use when user says \"生成图表\", \"画架构图\", \"AI绘图\", \"paper illustration\", \"generate diagram\", or needs visual figures for papers.

No reviews yet
0 installs
15 views
0.0% view→install

Install

$ agentstack add skill-raja21068-autoresearch-paper-illustration

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Paper Illustration? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Paper Illustration: Multi-Stage Claude-Supervised Figure Generation

Generate publication-quality illustrations using a multi-stage workflow with Claude as the STRICT supervisor/reviewer.

Core Design Philosophy

┌──────────────────────────────────────────────────────────────────────────┐
│                    MULTI-STAGE ITERATIVE WORKFLOW                        │
├──────────────────────────────────────────────────────────────────────────┤
│                                                                          │
│   User Request                                                           │
│       │                                                                  │
│       ▼                                                                  │
│   ┌─────────────┐                                                        │
│   │   Claude    │ ◄─── Step 1: Parse request, create initial prompt     │
│   │  (Planner)  │                                                        │
│   └──────┬──────┘                                                        │
│          │                                                               │
│          ▼                                                               │
│   ┌─────────────┐                                                        │
│   │   Gemini    │ ◄─── Step 2: Optimize layout description               │
│   │ (gemini-3-pro)│      - Refine component positioning                    │
│   │  Layout     │      - Optimize spacing and grouping                   │
│   └──────┬──────┘                                                        │
│          │                                                               │
│          ▼                                                               │
│   ┌─────────────┐                                                        │
│   │   Gemini    │ ◄─── Step 3: CVPR/NeurIPS style verification          │
│   │ (gemini-3-pro)│      - Check color palette compliance                  │
│   │  Style      │      - Verify arrow and font standards                 │
│   └──────┬──────┘                                                        │
│          │                                                               │
│          ▼                                                               │
│   ┌─────────────┐                                                        │
│   │ Paperbanana │ ◄─── Step 4: Render final image                       │
│   │ (gemini-3-  │      - High-quality image generation                   │
│   │ pro-image)  │      - Internal codename: Nano Banana Pro              │
│   └──────┬──────┘                                                        │
│          │                                                               │
│          ▼                                                               │
│   ┌─────────────┐                                                        │
│   │   Claude    │ ◄─── Step 5: STRICT visual review + SCORE (1-10)      │
│   │  (Reviewer) │      - Verify EVERY arrow direction                    │
│   │   STRICT!   │      - Verify EVERY block content                      │
│   └──────┬──────┘      - Verify aesthetics & visual appeal               │
│          │                                                               │
│          ▼                                                               │
│   Score ≥ 9? ──YES──► Accept & Output                                    │
│          │                                                               │
│          NO                                                              │
│          │                                                               │
│          ▼                                                               │
│   Generate SPECIFIC improvement feedback ──► Loop back to Step 2        │
│                                                                          │
└──────────────────────────────────────────────────────────────────────────┘

Constants

  • IMAGE_MODEL = gemini-3-pro-image-preview — Paperbanana (Nano Banana Pro) for image rendering
  • REASONING_MODEL = gemini-3-pro-preview — Gemini for layout optimization and style checking
  • MAX_ITERATIONS = 5 — Maximum refinement rounds
  • TARGET_SCORE = 9 — Minimum acceptable score (1-10) — RAISED FOR QUALITY
  • OUTPUT_DIR = figures/ai_generated/ — Output directory
  • APIKEYENV = GEMINI_API_KEY — Environment variable

Optional: Style reference (— style-ref: , opt-in)

Lets the user steer structural figure conventions (caption length, panel-count distribution, figure-to-table ratio in the parent paper) toward a reference paper. Default OFF — when the user does not pass — style-ref, do nothing differently from before.

Only when — style-ref: appears in $ARGUMENTS, run the helper FIRST, before generating prompts:

if [ ! -f tools/extract_paper_style.py ]; then
  echo "error: tools/extract_paper_style.py not found — re-run 'bash tools/install_aris.sh' to refresh the '.aris/tools' symlink (added in #174), or copy the helper manually from the ARIS repo" >&2
  exit 1
fi
CACHE=$(python3 tools/extract_paper_style.py --source "")
case $? in
  0) ;;                                       # use $CACHE/style_profile.md as structural guidance
  2) echo "warning: style-ref skipped (missing optional dep)" >&2 ;;
  3) echo "error: --style-ref source failed; aborting illustration" >&2 ; exit 1 ;;
  *) echo "error: helper failed unexpectedly; aborting illustration" >&2 ; exit 1 ;;
esac

Sources accepted: local TeX dir / file, local PDF, arXiv id, http(s) URL. Overleaf URLs/IDs are rejected — clone via /overleaf-sync setup first and pass the local clone path.

Strict rules (full contract in tools/extract_paper_style.py docstring):

  • Use style_profile.md to align caption length and figure density with the reference paper. The CVPR/ICLR/NeurIPS visual standards above still take precedence — --style-ref only refines length-and-density tendencies, never image content.
  • Never copy figure content, color palettes, or specific design elements from anything reachable through the cache. The visual design comes from the user's prompt, not the reference.
  • Never pass — style-ref (or the cache contents) to the Claude vision-checker / Gemini reasoning-checker sub-agents when they score the generated image — the image must be judged on its own merits.

CVPR/ICLR/NeurIPS Top-Tier Conference Style Guide

What "CVPR Style" Actually Means:

Visual Standards

  • Clean white background — No decorative patterns or gradients (unless subtle)
  • Sans-serif fonts — Arial, Helvetica, or Computer Modern; minimum 14pt
  • Subtle color palette — Not rainbow colors; use 3-5 coordinated colors
  • Print-friendly — Must be readable in grayscale (many reviewers print papers)
  • Professional borders — Thin (2-3px), solid colors, not flashy

Layout Standards

  • Horizontal flow — Left-to-right is the standard for pipelines
  • Clear grouping — Use subtle background boxes to group related modules
  • Consistent sizing — Similar components should have similar sizes
  • Balanced whitespace — Not cramped, not sparse

Arrow Standards (MOST CRITICAL)

  • Thick strokes — 4-6px minimum (thin arrows disappear when printed)
  • Clear arrowheads — Large, filled triangular heads
  • Dark colors — Black or dark gray (#333333); avoid colored arrows
  • Labeled — Every arrow should indicate what data flows through it
  • No crossings — Reorganize layout to avoid arrow crossings
  • CORRECT DIRECTION — Arrows must point to the RIGHT target!

Visual Appeal (科研风格 - Professional Academic Style)

目标:既不保守也不花哨,找到平衡点

✅ 应该有的视觉元素:
  • Subtle gradient fills — 淡雅的渐变填充(同色系从浅到深),不是炫彩
  • Rounded corners — 圆角矩形(6-10px radius),现代感但不夸张
  • Clear visual hierarchy — 通过大小、颜色深浅区分层次
  • Consistent color coding — 统一的配色方案(3-4种主色)
  • Internal structure — 大模块内部显示子组件(如Encoder内部的layer结构)
  • Professional typography — 清晰的标签,适当的字号层次
✅ 配色建议(学术专业):
  • Inputs: 柔和的绿色系 (#10B981 / #34D399)
  • Encoders: 专业的蓝色系 (#2563EB / #3B82F6)
  • Fusion: 优雅的紫色系 (#7C3AED / #8B5CF6)
  • Outputs: 温暖的橙色系 (#EA580C / #F97316)
  • Arrows: 黑色或深灰 (#333333 / #1F2937)
  • Background: 纯白 (#FFFFFF),不要花纹
❌ 要避免的过度装饰:
  • ❌ Rainbow color schemes (彩虹配色)
  • ❌ Heavy drop shadows (重阴影效果)
  • ❌ 3D effects / perspective (3D透视)
  • ❌ Excessive gradients (夸张的多色渐变)
  • ❌ Clip art / cartoon icons (卡通图标)
  • ❌ Decorative patterns in background (背景花纹)
  • ❌ Glowing effects (发光效果)
  • ❌ Too many small icons (过多小图标)
✓ 理想的视觉效果:
  • 一眼看上去专业、清晰
  • 适度的视觉吸引力,但不抢眼
  • 符合CVPR/NeurIPS论文的审美标准
  • 打印友好(灰度模式下也能清晰辨认)
  • 精心设计的学术图表,而不是PPT模板

What to AVOID (CRITICAL)

  • ❌ Rainbow color schemes (too many colors)
  • ❌ Thin, hairline arrows (arrows must be THICK)
  • ❌ Unlabeled connections
  • ❌ Plain boring rectangles (add some visual interest)
  • Over-decorated with shadows/glows/icons (too flashy)
  • ❌ Small text that's unreadable when printed
  • WRONG arrow directions — This is UNACCEPTABLE!

Scope

| Figure Type | Quality | Examples | |-------------|---------|----------| | Architecture diagrams | Excellent | Model architecture, pipeline, encoder-decoder | | Method illustrations | Excellent | Conceptual diagrams, algorithm flowcharts | | Conceptual figures | Good | Comparison diagrams, taxonomy trees |

Not for: Statistical plots (use /paper-figure), photo-realistic images

Workflow: MUST EXECUTE ALL STEPS

Step 0: Pre-flight Check

# Check API key
if [ -z "$GEMINI_API_KEY" ]; then
    echo "ERROR: GEMINI_API_KEY not set"
    echo "Get your key from: https://aistudio.google.com/app/apikey"
    echo "Set it: export GEMINI_API_KEY='your-key'"
    exit 1
fi

# Create output directory
mkdir -p figures/ai_generated

Step 1: Claude Plans the Figure (YOU ARE HERE)

CRITICAL: Claude must first analyze the user's request and create a detailed prompt.

Parse the input: $ARGUMENTS

Claude's task:

  1. Understand what figure the user wants
  2. Identify all components, connections, data flow
  3. Create a detailed, structured prompt for Gemini
  4. Include style requirements AND visual appeal requirements

Prompt Template for Claude to generate:

Create a PROFESSIONAL, VISUALLY APPEALING publication-quality academic diagram following CVPR/ICLR/NeurIPS standards.

## Visual Style: 科研风格 (Academic Professional Style)
### 目标:平衡 — 既不保守也不花哨

#### DO (应该有):
- **Subtle gradients** — 同色系淡雅渐变(如 #2563EB → #3B82F6),不是多色炫彩
- **Rounded corners** — 圆角矩形(6-10px),现代感
- **Clear visual hierarchy** — 通过大小、深浅区分层次
- **Internal structure** — 大模块内显示子组件结构
- **Consistent color coding** — 统一的3-4色方案
- **Professional polish** — 精致但不夸张

#### DON'T (不要有):
- ❌ Rainbow/multi-color gradients (彩虹渐变)
- ❌ Heavy drop shadows (重阴影)
- ❌ 3D effects / perspective (3D效果)
- ❌ Glowing effects (发光效果)
- ❌ Excessive decorative icons (过多装饰图标)
- ❌ Plain boring rectangles (完全平淡的方块)

#### 理想效果:
像顶会论文中精心设计的架构图 — 专业、清晰、有适度的视觉吸引力

## Figure Type
[Architecture Diagram / Pipeline / Comparison / etc.]

## Components to Include (BE SPECIFIC ABOUT CONTENT)
1. [Component 1]:
   - Label: "[exact text]"
   - Sub-label: "[smaller text below]"
   - Position: [left/center/right, top/middle/bottom]
   - Style: [border color, fill, internal structure]
2. [Component 2]: ...

## Layout
- Direction: [left-to-right / top-to-bottom]
- Spacing: [tight / normal / loose]
- Grouping: [how components should be grouped]

## Connections (BE EXPLICIT ABOUT DIRECTION)
EXACT arrow specifications:
1. [Component A] → [Component B]: Arrow goes FROM A TO B, label it "[data type]"
2. [Component C] → [Component D]: Arrow goes FROM C TO D, label it "[data type]"
...
VERIFY: Each arrow must point to the CORRECT target!

## Style Requirements (CVPR/ICLR/NeurIPS Standard)

### Visual Style
- Color palette: Professional academic colors
  - Inputs: Green (#10B981)
  - Encoders: Blue (#2563EB)
  - Fusion modules: Purple (#7C3AED)
  - Outputs: Orange (#EA580C)
- Font: Sans-serif (Arial/Helvetica), minimum 14pt, bold for labels
- Background: Clean white, no patterns
- Blocks: Rounded rectangles (8-12px radius), subtle gradient fill, colored border (2-3px)
- Subtle shadows for depth effect
- Print-friendly (must work in grayscale)

### CRITICAL: Arrow & Data Flow Requirements
1. **ALL arrows must be VERY THICK** - minimum 5-6px stroke width
2. **ALL arrows must have CLEAR arrowheads** - large, visible triangular heads
3. **ALL arrows must be BLACK or DARK GRAY** - not colored
4. **Label EVERY arrow** with what data flows through it
5. **VERIFY arrow direction** - each arrow MUST point to the correct target
6. **No ambiguous connections** - every arrow should have a clear source and destination

### Logic Clarity Requirements
1. **Data flow must be immediately obvious** - viewer should understand the pipeline in 5 seconds
2. **No crossing arrows** - reorganize layout to avoid arrow crossings
3. **Consistent direction** - maintain left-to-right or top-to-bottom flow throughout
4. **Group related components** - use subtle background boxes or spacing to group modules
5. **Clear hierarchy** - main components larger, sub-components smaller

## Additional Requirements
[Any specific requirements from user]

Step 2: Gemini Layout Optimization (gemini-3-pro)

Claude sends the initial prompt to Gemini (gemini-3-pro) for layout optimization.

#!/bin/bash
# Step 2: Optimize layout using Gemini gemini-3-pro
# This step refines component positioning and spacing

set -e

OUTPUT_DIR="figures/ai_generated"
mkdir -p "$OUTPUT_DIR"

API_KEY="${GEMINI_API_KEY}"
URL="https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-preview:generateContent?key=$API_KEY"

# The initial prompt from Claude
INITIAL_PROMPT='[Claude fills in the detailed prompt here]'

# Layout optimization request
LAYOUT_REQUEST="You are an expert in academic figure layout design for CVPR/NeurIPS papers.

Analyze this figure request and provide an OPTIMIZED LAYOUT DESCRIPTION:

$INITIAL_PROMPT

Provide:
1. **Optimized Component Positions**: Exact positions (left/center/right, top/middle/bottom) for each component
2. **Spacing Recommendations**: Specific spacing between components
3. **Grouping Strategy**: Which components should be visually grouped together
4. **Arrow Routing**: Optimal paths for arrows to avoid crossings
5. **Visual Hierarchy**: Size recommendations for main vs sub-components

Output a DETAILED layout specification that will be used for rendering."

# Build JSON payload
python3  "$OUTPUT_DIR/layout_description.txt"

Step 3: Gemini Style Verification (gemini-3-pro)

Claude sends the optimized layout to Gemini for CVPR/NeurIPS style verification.

#!/bin/bash
# Step 3: Verify and enhance style compliance using Gemini gemini-3-pro

API_KEY="${GEMINI_API_KEY}"
URL="https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-preview:generateContent?key=$API_KEY"

# Read layout from previous step
LAYOUT=$(cat figures/ai_generated/layout_description.txt)

# Style verification request
STYLE_REQUEST="You are a CVPR/NeurIPS paper figure reviewer specializing in visual standards.

Review and ENHANCE this figure specification for top-tier conference compliance:

$LAYOUT

Ensure compliance with:
1. **Color Palette**: Use professional academic colors (green for inputs, blue for encoders, purple for fusion, orange for outputs)
2. **Arrow Standards**: Thick (5-6px), black/dark gray, clear arrowheads, all labeled
3. **Font Standards**: Sans-serif, minimum 14pt, readable in print
4. **Visual Appeal (科研风格)**:
   - ✅ Subtle same-color gradients, rounded corners (6-10px), internal structure visible
   - ❌ NO heavy shadows, NO glowing effects, NO rainbow gradients

Output an ENHANCED figure specification with explicit style instructions for rendering."

# Build

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [raja21068](https://github.com/raja21068)
- **Source:** [raja21068/AutoResearch](https://github.com/raja21068/AutoResearch)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.