AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Arabench

skill-moshe-ship-hurmoz-arabench · by Moshe-ship

Arabic LLM benchmarking across 8 quality categories

No reviews yet
0 installs
23 views
0.0% view→install

Install

$ agentstack add skill-moshe-ship-hurmoz-arabench

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-moshe-ship-hurmoz-arabench)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Arabench? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

arabench — معيار جودة العربية للنماذج

أداة لتقييم جودة الذكاء الاصطناعي بالعربي عبر 8 فئات.

الأوامر

تشغيل المعيار الكامل

arabench run

النتيجة المتوقعة:

╔══════════════════════════════════════════════╗
║          arabench — Full Benchmark           ║
╠════════════════╦═══════╦═══════╦═════════════╣
║ Category       ║ GPT-4 ║ Claude║ Gemini      ║
╠════════════════╬═══════╬═══════╬═════════════╣
║ Translation    ║  82   ║  87   ║  79         ║
║ Grammar        ║  76   ║  84   ║  74         ║
║ Dialect        ║  68   ║  71   ║  65         ║
║ Diacritization ║  54   ║  62   ║  51         ║
║ Summarization  ║  78   ║  83   ║  76         ║
║ QA             ║  81   ║  85   ║  78         ║
║ Generation     ║  75   ║  80   ║  73         ║
║ Culture        ║  70   ║  77   ║  66         ║
╠════════════════╬═══════╬═══════╬═════════════╣
║ OVERALL        ║  73.0 ║  78.6 ║  70.3       ║
╚════════════════╩═══════╩═══════╩═════════════╝

اختبار سريع لمزوّد واحد

arabench quick claude
arabench quick gpt4
arabench quick gemini

النتيجة — نتائج سريعة خلال 30 ثانية لأهم 3 فئات (translation, grammar, qa).

مقارنة مزوّدين

arabench compare claude gpt4

النتيجة المتوقعة:

claude vs gpt4 — Arabic Quality Comparison
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Translation:    claude 87 ████████▋  vs  gpt4 82 ████████▏
Grammar:        claude 84 ████████▍  vs  gpt4 76 ███████▌
Dialect:        claude 71 ███████    vs  gpt4 68 ██████▊
Winner: claude (+5.6 avg)

عرض لوحة النتائج

arabench leaderboard

يعرض ترتيب جميع المزوّدين المُقيّمين حسب المعدل العام.

شرح فئة تقييم

arabench explain translation
arabench explain dialect
arabench explain diacritization

الفئات الثمانية بالتفصيل

| الفئة | الوصف | أمثلة الاختبار | |-------|-------|----------------| | translation | دقة الترجمة عربي↔إنجليزي | مصطلحات تقنية، تعابير اصطلاحية | | grammar | صحة القواعد النحوية والصرفية | إعراب، تصريف أفعال، جمع تكسير | | dialect | فهم اللهجات الخمس | مصري، خليجي، شامي، مغاربي، عراقي | | diacritization | دقة التشكيل | نصوص بدون تشكيل يُطلب تشكيلها | | summarization | جودة التلخيص العربي | مقالات إخبارية، نصوص أكاديمية | | qa | الإجابة على أسئلة بالعربي | ثقافية، تاريخية، علمية | | generation | جودة توليد النصوص | مقالات، قصص، محتوى تسويقي | | culture | الوعي الثقافي العربي | عادات، أمثال، سياق اجتماعي |

نظام التقييم

  • 90-100: ممتاز — أداء يقارب المتحدث الأصلي
  • 80-89: جيد جدا — أخطاء نادرة وطفيفة
  • 70-79: جيد — يفهم السياق لكن فيه أخطاء ملحوظة
  • 60-69: مقبول — يحتاج تدقيق بشري
  • أقل من 60: ضعيف — غير موثوق للاستخدام الإنتاجي

متى تستخدم

  • المستخدم يسأل "أي نموذج أفضل بالعربي؟"
  • يريد مقارنة بين Claude و GPT أو غيرهم
  • يريد يعرف نقاط ضعف نموذج معين بالعربي
  • يختار نموذج لمشروع يتطلب عربي عالي الجودة
  • يريد يفهم ليش نموذج معين ضعيف بالتشكيل أو اللهجات

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.