AgentStack
SKILL verified MIT Self-run

Asyncopenai Concurrency Httpx Pool

skill-kennethkhoocy-applied-micro-skills-asyncopenai-concurrency-httpx-pool · by kennethkhoocy

|

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add skill-kennethkhoocy-applied-micro-skills-asyncopenai-concurrency-httpx-pool

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Asyncopenai Concurrency Httpx Pool? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

AsyncOpenAI Concurrency: the Hidden httpx Pool Cap

Problem

Async batch scorers typically gate concurrency with asyncio.Semaphore(N). Raising N above ~100 silently does nothing: the OpenAI SDK's default httpx transport caps the connection pool at max_connections=100, so excess tasks queue inside httpx instead of reaching the provider. The semaphore looks like the throttle but is not the binding constraint — there is no error, just a throughput ceiling.

Context / Trigger Conditions

  • asyncio.Semaphore(N) with N > 100 around client.chat.completions.create

shows the same throughput as N = 100

  • Client constructed as AsyncOpenAI(api_key=..., base_url=...) with no

http_client argument (the default transport)

  • Provider is known to allow high concurrency (DeepSeek v4-flash: ~2500)
  • Symptom check: requests-in-flight measured at the server never exceeds ~100

Solution

Size the httpx pool to the semaphore when constructing the client:

import httpx
from openai import AsyncOpenAI

CONCURRENCY = 2000
client = AsyncOpenAI(
    api_key=..., base_url="https://api.deepseek.com",
    http_client=httpx.AsyncClient(limits=httpx.Limits(
        max_connections=CONCURRENCY,
        max_keepalive_connections=CONCURRENCY)))
sem = asyncio.Semaphore(CONCURRENCY)

Both edits are required; either alone caps the other. Keep the per-request retry loop — at high concurrency transient failures are more likely, and the retry envelope is what turns them into non-events.

Verification

Throughput scales with N. Verified 2026-07-17 on DeepSeek v4-flash (deepseek-chat, JSON-mode unit scoring, ~1.5k-token prompts): at semaphore 50 a cold 13.8k-request chunk took ~70 min; at semaphore 2000 + matched pool, a 14.4k-request chunk took ~11 min (~40 req/s sustained, ~250/s burst on a 1.7k-request tail chunk), 0 failed requests, 0 schema-invalid responses. Effective speedup ~7x rather than 40x — server-side queuing absorbs the rest — but with zero reliability cost.

Example

Specialist Directors US T1 campaign: final 8 chunks (100,243 calls) completed in ~55 minutes for $8.08 after the fix, versus a projected ~9 hours at the old setting. The edit is two lines in the scorer; the semaphore constant alone would have been a silent no-op.

Notes

  • DeepSeek publishes no hard rate limit and handled 2000 in-flight cleanly;

the practical ceiling reported is ~2500. Other providers enforce RPM/TPM caps — check before sizing.

  • Windows: default asyncio proactor loop handled 2000 sockets without

tuning; no ulimit-style adjustment needed.

  • Companion ops lesson from the same campaign: when a run's plan changes

scale (two-night legs -> one-shot), re-audit the launch glue's hardcoded limits — a wrapper MAX-HOURS=10 safety cap sized for the old plan hard-killed a healthy runner at 97/108 chunks. Caps and budgets in supervisor scripts must be revisited whenever expected duration changes.

  • Concurrency is a pure throughput knob: per-request outputs are unchanged

(temperature 0, independent requests), so raising it mid-campaign does not create a scoring seam — unlike model/prompt changes, which do (see [llm-campaign-drift-gate]).

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.