The scoreboard, publicلوحة النتائج، علنية
Benchmarkالقياس
Quality claims are measurements or they are marketing. Two honest tiers below: the offline gates block every deploy when they regress; the live diagnostics are measured against the production catalog and published as they are — including the ones we currently fail.ادعاءات الجودة إما قياسات وإما تسويق. طبقتان صريحتان أدناه: البوابات دون اتصال توقف كل نشر عند أي تراجع؛ أما القياسات الحية فتُقاس على فهرس الإنتاج وتُنشَر كما هي — بما فيها ما نفشل فيه حاليًا.
Measuredقيس في 2026-07-29 03:41 UTC
Blocking gates (offline — every deploy)بوابات موقِفة (دون اتصال — كل نشر)
- PASSArabic normalization (known-hard cases)33 cases
- PASSDuplicate detection (incl. trap non-dupes)10 pairs (incl. traps)
- PASSPerson duplicate detection (incl. traps)16 pairs (incl. traps)
- PASSPrompt-injection defense patterns8 patterns, boundary invariants hold
- PASSInfobox cast parsing (precision-first)16 fixtures, 48 checks
- PASSCorroboration gate (promotion policy)22 cases (promote + must-not-promote)
- PASSBroadcast claim gate (clean + poisoned)8 cases (clean + poisoned)
Live diagnostics (measured on the production catalog, failures included)قياسات حية (على فهرس الإنتاج، بما فيها الإخفاقات)
- PASSSearch recall (variant golden set)top-3 recall 100.0% over 117 gating cases (target 97%)
- PASSsearch-latencyp50 0ms · p95 246ms over 24 runs (budget 1200ms)
- PASSPhonetic Latin–Arabic recallrecall 100% (8/8)
- PASSPlot-description search (semantic)top-5 recall 80.0% (target 60%) | misses: "daily life in a damascus alley neighbourhood under the french mandate" → wanted bab-al-hara ; "pre-islamic tribal war epic of revenge sparked by the killing of a king" → wanted al-zir-salem ; "saudi drama about a riyadh neighbourhood in the 1970s" → wanted al-asouf
- PASSShowcase demos — every /try button, verified live11/11 showcased demos resolve top-3 live
- PASSnormalize-sql-drift30 golden inputs identical in both homes
- PASSNatural-language query parsingaccuracy 100.0% (target 75%)
- PASSAI extraction accuracy5/5 cases (cost ~$0.0752)not re-scored sinceلم يُعَد قياسه منذ 2026-07-26
- PASSAsk MSDB (citation-enforced answers)25 questions, 20 answered, 100% citednot re-scored sinceلم يُعَد قياسه منذ 2026-07-26
- PASSName transliteration accuracy9/10 names | misses: reverse: unknown Egyptian romanization → null@0.3not re-scored sinceلم يُعَد قياسه منذ 2026-07-26
- PASSDiscovery triage accuracy10/10 sightings
- PASSFilmography resolution (north star)resolution 98.8% (2462/2491 works across 102 pros; target 85%) | تيم حسن 10/10 · سلوم حداد 16/16 · جمال سليمان 9/9 · دريد لحام 11/11 · نادين لبكي 13/13 · زياد دويري 8/8 · حاتم علي 20/20 · هاني أبو أسعد 11/12not re-scored sinceلم يُعَد قياسه منذ 2026-07-26
- SKIPFreshness latency ledger (advisory)ADVISORY (not blocking): 11 discovery-latency samples (< 20 floor); 114 benchmark rows show first-to-list 74% over BACKFILLED titles (not the discovery-originated population the metric targets). Pollers run on the fra1 cron; the gate arms as discovery drafts accrue. Targets: median <72h, p90 <168h, first-to-list ≥90%.
SKIP rows are suites awaiting a live dependency (API key or database) — reported as skipped, never counted as passes. FAIL rows stay published; features they gate stay off the site until the number clears the bar.صفوف SKIP مجموعات بانتظار اعتمادية حية (مفتاح واجهة أو قاعدة بيانات) — تُبلَّغ متخطّاة ولا تُحتسب نجاحًا أبدًا. وصفوف FAIL تبقى منشورة؛ والميزات المرهونة بها تبقى موقوفة عن الموقع حتى يجتاز الرقم العتبة.
The comparison protocol (vs IMDb-class databases)بروتوكول المقارنة (مقابل قواعد بيانات من فئة IMDb)
Thirty real queries — ten chat-Arabic, ten dialect/typo, ten plot descriptions — run manually by a humanagainst other databases and recorded, then run automatically against this engine by the eval harness. Repeated quarterly. We sample; we never crawl or scrape anyone's catalog.
ثلاثون استعلامًا حقيقيًا — عشرة بعربية الدردشة، وعشرة بالعامية أو الأخطاء المطبعية، وعشرة بوصف الحبكة — تُشغَّل يدويًا بواسطة إنسان مقابل قواعد البيانات الأخرى وتُسجَّل، ثم تُشغَّل آليًا مقابل هذا المحرّك عبر منظومة التقييم. وتتكرّر كل ربع سنة. نحن نأخذ عيّنات؛ ولا نزحف أو ننسخ فهرس أحد إطلاقًا.
First recorded sampling run: pending — results publish here, including the queries we lose. Until then, run the impossible searches yourself.
أول تشغيل عيّنات مُسجَّل: قيد الانتظار — تُنشَر النتائج هنا، بما فيها الاستعلامات التي نخسرها. وحتى ذلك الحين، جرّب عمليات البحث المستحيلة بنفسك.
Method details on how we know. Latency budgets (search p95 < 300ms, suggestions < 120ms) publish here once measured against the production deployment.
تفاصيل الطريقة على كيف نعرف. وتُنشَر ميزانيات الكمون (بحث p95 < ٣٠٠ مللي ثانية، اقتراحات < ١٢٠ مللي ثانية) هنا بمجرّد قياسها مقابل نشر الإنتاج.