AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Apache-2.0 Self-run

Gemini Count In Video

skill-benchflow-ai-skillsbench-gemini-count-in-video · by benchflow-ai

Analyze and count objects in videos using Google Gemini API (object counting, pedestrian detection, vehicle tracking, and surveillance video analysis).

No reviews yet
0 installs
3 views
0.0% view→install

Install

$ agentstack add skill-benchflow-ai-skillsbench-gemini-count-in-video

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-benchflow-ai-skillsbench-gemini-count-in-video)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Gemini Count In Video? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Gemini Video Understanding Skill

Purpose

This skill enables video analysis and object counting using the Google Gemini API, with a focus on counting pedestrians, detecting objects, tracking movement, and analyzing surveillance footage. It supports precise prompting for differentiated counting (e.g., pedestrians vs cyclists vs vehicles).

When to Use

  • Counting pedestrians, vehicles, or other objects in surveillance videos
  • Distinguishing between different types of objects (walkers vs cyclists, cars vs trucks)
  • Analyzing traffic patterns and movement through a scene
  • Processing multiple videos for batch object counting
  • Extracting structured count data from video footage

Required Libraries

The following Python libraries are required:

from google import genai
from google.genai import types
import os
import time

Input Requirements

  • File formats: MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
  • Size constraints:
  • Use inline bytes for small files (rule of thumb: 20MB)

myfile = client.files.upload(file="surveillance.mp4")

Wait for processing

while myfile.state.name == "PROCESSING": time.sleep(5) myfile = client.files.get(name=myfile.name)

if myfile.state.name == "FAILED": raise ValueError("Video processing failed")

Prompt for counting pedestrians with clear exclusion criteria

prompt = """Count the total number of pedestrians who are WALKING through the scene in this surveillance video.

IMPORTANT RULES:

  • ONLY count people who are walking on foot
  • DO NOT count people riding bicycles
  • DO NOT count people driving cars or other vehicles
  • Count each unique pedestrian only once, even if they appear in multiple frames

Provide your answer as a single integer number representing the total count of pedestrians. Answer with just the number, nothing else. Your answer should be enclosed in and tags, such as 5. """

response = client.models.generate_content( model="gemini-2.0-flash-exp", contents=[prompt, myfile], )

Parse the response

responsetext = response.text.strip() match = re.search(r"(\d+)", responsetext) if match: count = int(match.group(1)) print(f"Pedestrian count: {count}") else: print("Could not parse count from response")


### Batch Processing Multiple Videos

```python
from google import genai
import os
import time
import re

def upload_and_wait(client, file_path: str, max_wait_s: int = 300):
    """Upload video and wait for processing."""
    myfile = client.files.upload(file=file_path)
    waited = 0
    
    while myfile.state.name == "PROCESSING" and waited 
Cyclists: 
Vehicles: 
"""

response = client.models.generate_content(
    model="gemini-2.0-flash-exp",
    contents=[prompt, myfile],
)

# Parse multiple counts
text = response.text
pedestrians = int(re.search(r'Pedestrians:\s*(\d+)', text).group(1))
cyclists = int(re.search(r'Cyclists:\s*(\d+)', text).group(1))
vehicles = int(re.search(r'Vehicles:\s*(\d+)', text).group(1))

Using Answer Tags for Reliable Parsing

# Request structured output with XML-like tags
prompt = """Count the total number of pedestrians walking through the scene.

You should reason and think step by step. Provide your answer as a single integer.
Your answer should be enclosed in  and  tags, such as 5.
"""

response = client.models.generate_content(
    model="gemini-2.0-flash-exp",
    contents=[prompt, myfile],
)

# Robust extraction
match = re.search(r"(\d+)", response.text)
if match:
    count = int(match.group(1))
else:
    # Fallback: try to find any number in response
    numbers = re.findall(r'\d+', response.text)
    count = int(numbers[0]) if numbers else 0

Best Practices

  • Use the File API for all surveillance videos (typically >20MB) and always wait for processing to complete.
  • Be specific in prompts: Clearly define what to count and what to exclude (e.g., "walking pedestrians only, not cyclists").
  • Use structured output formats: Request answers in specific formats (like N) for reliable parsing.
  • Ask for reasoning: Include "think step by step" to improve counting accuracy.
  • Handle edge cases: Specify rules for partial appearances, people entering/exiting frame, and mode changes.
  • Use gemini-2.0-flash-exp or gemini-2.5-flash: These models provide good balance of speed and accuracy for object counting.
  • Test with sample videos: Verify prompt effectiveness on representative samples before batch processing.

Error Handling

import time

def upload_and_wait(client, file_path: str, max_wait_s: int = 300):
    """Upload video and wait for processing with timeout."""
    myfile = client.files.upload(file=file_path)
    waited = 0

    while myfile.state.name == "PROCESSING" and waited  tags."""
        
        response = client.models.generate_content(
            model="gemini-2.0-flash-exp",
            contents=[prompt, myfile],
        )
        
        # Try structured parsing first
        match = re.search(r"(\d+)", response.text)
        if match:
            return int(match.group(1))
        
        # Fallback to any number found
        numbers = re.findall(r'\d+', response.text)
        if numbers:
            return int(numbers[0])
        
        print(f"Warning: Could not parse count, defaulting to 0")
        return 0
        
    except Exception as e:
        print(f"Error processing video: {e}")
        return 0

Common issues:

  • Upload processing stuck: Use timeout logic and fail gracefully after max wait time
  • Ambiguous responses: Use structured output tags like `` for reliable parsing
  • Rate limits: Add retry logic with exponential backoff for batch processing
  • Inconsistent counts: Be very explicit in prompts about counting rules and exclusions

Limitations

  • Counting accuracy depends on video quality, camera angle, and object size/distance
  • Very crowded scenes may have higher counting variance
  • Occlusion (objects blocking each other) can affect accuracy
  • Long videos require longer processing times (typically 5-30 seconds per video)
  • The model may occasionally misclassify similar objects (e.g., motorcyclist as cyclist)
  • For highest accuracy, use clear prompts with explicit inclusion/exclusion criteria

Version History

  • 1.0.0 (2026-01-21): Tailored for pedestrian traffic counting with focus on object counting, differentiation, and batch processing

Resources

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.