AgentStack
SKILL verified MIT Self-run

Computer Vision Pipeline

skill-curiositech-some-claude-skills-computer-vision-pipeline · by curiositech

Build production computer vision pipelines for object detection, tracking, and video analysis. Handles drone footage, wildlife monitoring, and real-time detection. Supports YOLO, Detectron2,

No reviews yet
0 installs
10 views
0.0% view→install

Install

$ agentstack add skill-curiositech-some-claude-skills-computer-vision-pipeline

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Computer Vision Pipeline? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Computer Vision Pipeline

Expert in building production-ready computer vision systems for object detection, tracking, and video analysis.

When to Use

Use for:

  • Drone footage analysis (archaeological surveys, conservation)
  • Wildlife monitoring and tracking
  • Real-time object detection systems
  • Video preprocessing and analysis
  • Custom model training and inference
  • Multi-object tracking (MOT)

NOT for:

  • Simple image filters (use Pillow/PIL)
  • Photo editing (use Photoshop/GIMP)
  • Face recognition APIs (use AWS Rekognition)
  • Basic OCR (use Tesseract)

Technology Selection

Object Detection Models

| Model | Speed (FPS) | Accuracy (mAP) | Use Case | |-------|-------------|----------------|----------| | YOLOv8 | 140 | 53.9% | Real-time detection | | Detectron2 | 25 | 58.7% | High accuracy, research | | EfficientDet | 35 | 55.1% | Mobile deployment | | Faster R-CNN | 10 | 42.0% | Legacy systems |

Timeline:

  • 2015: Faster R-CNN (two-stage detection)
  • 2016: YOLO v1 (one-stage, real-time)
  • 2020: YOLOv5 (PyTorch, production-ready)
  • 2023: YOLOv8 (state-of-the-art)
  • 2024: YOLOv8 is industry standard for real-time

Decision tree:

Need real-time (>30 FPS)? → YOLOv8
Need highest accuracy? → Detectron2 Mask R-CNN
Need mobile deployment? → YOLOv8-nano or EfficientDet
Need instance segmentation? → Detectron2 or YOLOv8-seg
Need custom objects? → Fine-tune YOLOv8

Common Anti-Patterns

Anti-Pattern 1: Not Preprocessing Frames Before Detection

Novice thinking: "Just run detection on raw video frames"

Problem: Poor detection accuracy, wasted GPU cycles.

Wrong approach:

# ❌ No preprocessing - poor results
import cv2
from ultralytics import YOLO

model = YOLO('yolov8n.pt')
video = cv2.VideoCapture('drone_footage.mp4')

while True:
    ret, frame = video.read()
    if not ret:
        break

    # Raw frame detection - no normalization, no resizing
    results = model(frame)
    # Poor accuracy, slow inference

Why wrong:

  • Video resolution too high (4K = 8.3 megapixels per frame)
  • No normalization (pixel values 0-255 instead of 0-1)
  • Aspect ratio not maintained
  • GPU memory overflow on high-res frames

Correct approach:

# ✅ Proper preprocessing pipeline
import cv2
import numpy as np
from ultralytics import YOLO

model = YOLO('yolov8n.pt')
video = cv2.VideoCapture('drone_footage.mp4')

# Model expects 640x640 input
TARGET_SIZE = 640

def preprocess_frame(frame):
    # Resize while maintaining aspect ratio
    h, w = frame.shape[:2]
    scale = TARGET_SIZE / max(h, w)
    new_w, new_h = int(w * scale), int(h * scale)

    resized = cv2.resize(frame, (new_w, new_h), interpolation=cv2.INTER_LINEAR)

    # Pad to square
    pad_w = (TARGET_SIZE - new_w) // 2
    pad_h = (TARGET_SIZE - new_h) // 2

    padded = cv2.copyMakeBorder(
        resized,
        pad_h, TARGET_SIZE - new_h - pad_h,
        pad_w, TARGET_SIZE - new_w - pad_w,
        cv2.BORDER_CONSTANT,
        value=(114, 114, 114)  # Gray padding
    )

    # Normalize to 0-1 (if model expects it)
    # normalized = padded.astype(np.float32) / 255.0

    return padded, scale

while True:
    ret, frame = video.read()
    if not ret:
        break

    preprocessed, scale = preprocess_frame(frame)
    results = model(preprocessed)

    # Scale bounding boxes back to original coordinates
    for box in results[0].boxes:
        x1, y1, x2, y2 = box.xyxy[0]
        x1, y1, x2, y2 = x1/scale, y1/scale, x2/scale, y2/scale

Performance comparison:

  • Raw 4K frames: 5 FPS, 72% mAP
  • Preprocessed 640x640: 45 FPS, 89% mAP

Timeline context:

  • 2015: Manual preprocessing required
  • 2020: YOLOv5 added auto-resize
  • 2023: YOLOv8 has smart preprocessing but explicit control is better

Anti-Pattern 2: Processing Every Frame in Video

Novice thinking: "Run detection on every single frame"

Problem: 99% of frames are redundant, wasting compute.

Wrong approach:

# ❌ Process every frame (30 FPS video = 1800 frames/min)
import cv2
from ultralytics import YOLO

model = YOLO('yolov8n.pt')
video = cv2.VideoCapture('drone_footage.mp4')

detections = []

while True:
    ret, frame = video.read()
    if not ret:
        break

    # Run detection on EVERY frame
    results = model(frame)
    detections.append(results)

# 10-minute video = 18,000 inferences (15 minutes on GPU)

Why wrong:

  • Adjacent frames are nearly identical
  • Wasting 95% of compute on duplicate work
  • Slow processing time
  • Massive storage for results

Correct approach 1: Frame sampling

# ✅ Sample every Nth frame
import cv2
from ultralytics import YOLO

model = YOLO('yolov8n.pt')
video = cv2.VideoCapture('drone_footage.mp4')

SAMPLE_RATE = 30  # Process 1 frame per second (if 30 FPS video)

frame_count = 0
detections = []

while True:
    ret, frame = video.read()
    if not ret:
        break

    frame_count += 1

    # Only process every 30th frame
    if frame_count % SAMPLE_RATE == 0:
        results = model(frame)
        detections.append({
            'frame': frame_count,
            'timestamp': frame_count / 30.0,
            'results': results
        })

# 10-minute video = 600 inferences (30 seconds on GPU)

Correct approach 2: Adaptive sampling with scene change detection

# ✅ Only process when scene changes significantly
import cv2
import numpy as np
from ultralytics import YOLO

model = YOLO('yolov8n.pt')
video = cv2.VideoCapture('drone_footage.mp4')

def scene_changed(prev_frame, curr_frame, threshold=0.3):
    """Detect scene change using histogram comparison"""
    if prev_frame is None:
        return True

    # Convert to grayscale
    prev_gray = cv2.cvtColor(prev_frame, cv2.COLOR_BGR2GRAY)
    curr_gray = cv2.cvtColor(curr_frame, cv2.COLOR_BGR2GRAY)

    # Calculate histograms
    prev_hist = cv2.calcHist([prev_gray], [0], None, [256], [0, 256])
    curr_hist = cv2.calcHist([curr_gray], [0], None, [256], [0, 256])

    # Compare histograms
    correlation = cv2.compareHist(prev_hist, curr_hist, cv2.HISTCMP_CORREL)

    return correlation  30:  # Only long tracks
        print(f"Dolphin {track_id} appeared in {len(trajectory)} frames")
        # Calculate movement, speed, etc.

Tracking benefits:

  • Count unique objects (not just detections per frame)
  • Build trajectories and movement patterns
  • Analyze behavior over time
  • Filter out brief false positives

Tracking algorithms: | Algorithm | Speed | Robustness | Occlusion Handling | |-----------|-------|------------|---------------------| | ByteTrack | Fast | Good | Excellent | | SORT | Very Fast | Fair | Fair | | DeepSORT | Medium | Excellent | Good | | BotSORT | Medium | Excellent | Excellent |


Production Checklist

□ Preprocess frames (resize, pad, normalize)
□ Sample frames intelligently (1 FPS or scene change detection)
□ Use batch inference (16-32 images per batch)
□ Tune NMS thresholds for your use case
□ Implement tracking if analyzing video
□ Log inference time and GPU utilization
□ Handle edge cases (empty frames, corrupted video)
□ Save results in structured format (JSON, CSV)
□ Visualize detections for debugging
□ Benchmark on representative data

When to Use vs Avoid

| Scenario | Appropriate? | |----------|--------------| | Analyze drone footage for archaeology | ✅ Yes - custom object detection | | Track wildlife in video | ✅ Yes - detection + tracking | | Count people in crowd | ✅ Yes - dense object detection | | Real-time security camera | ✅ Yes - YOLOv8 real-time | | Filter vacation photos | ❌ No - use photo management apps | | Face recognition login | ❌ No - use AWS Rekognition API | | Read license plates | ❌ No - use specialized OCR |


References

  • /references/yolo-guide.md - YOLOv8 setup, training, inference patterns
  • /references/video-processing.md - Frame extraction, scene detection, optimization
  • /references/tracking-algorithms.md - ByteTrack, SORT, DeepSORT comparison

Scripts

  • scripts/video_analyzer.py - Extract frames, run detection, generate timeline
  • scripts/model_trainer.py - Fine-tune YOLO on custom dataset, export weights

This skill guides: Computer vision | Object detection | Video analysis | YOLO | Tracking | Drone footage | Wildlife monitoring

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.