We gave the Consilium panel the most famous line in Arabic poetry — the opening of Imru' al-Qays's mu'allaqa — and asked for a word-by-word grammatical parse. Five models answered in Arabic independently, graded each other blind, and the verdict was fact-checked on the live web. The result: even the winner made mistakes — and the fact-check corrected the lead model itself.
أعرِب إعرابًا مفصّلًا صدرَ معلّقة امرئ القيس: «قِفا نَبْكِ مِن ذِكرى حَبيبٍ ومَنزِلِ» كلمةً كلمةً، مع بيان سبب جزم الفعل «نَبْكِ» وعلامة جزمه.
Won the blind ranking with a structured, readable parse — yet it erred too: it called «قِفا» built on sukūn (the correct answer is deletion of the nūn, dual address), and named the kasra as the jussive marker instead of the deleted weak letter.
The only answer matching the grammar books on every point: «قِفا» built on nūn-deletion with the dual alif as subject, «نَبْكِ» jussive as the answer to a request, marked by weak-letter deletion — citing mu'allaqa commentaries live while answering.
A solidly-structured parse that fell into the classic grammar-student error: it claimed the jussive marker is sukūn, insisted the verb is not weak-ending, and explained the kasra as avoiding two sukūns meeting.
Got the reason right (answer to the imperative) but never committed to the marker: it offered three competing answers — nūn-deletion as "the best known", sukūn, and weak-letter deletion — and left the reader to choose.
Fell apart in formal Arabic: Chinese and Russian words leaked into its parse mid-sentence, and it analyzed «نَبْكِ» as one of the five verbs jussive by nūn-deletion — which it is not.
The lead model carried its error into the draft verdict: "the marker is sukūn; the verb is sound-ending." Then the fact-check moved: Phase ④ took every claim to the live web, found the mu'allaqa commentaries and jussive rules stating the opposite — and marked the lead's claim as contradicted by the sources, flagging the verdict.
Phase ⑤ rewrote the verdict on the sources, not the lead's opinion: the jussive marker is deletion of the weak letter, and the kasra is the trace pointing to the deleted yā'. The model that drafted the verdict is the same one that got corrected — documented in the change log below.
قِفا: فعل أمر مبني على حذف النون، والألف ضمير متصل مبني على السكون في محل رفع فاعل، والخطاب للمثنى [9][13].
نَبْكِ: فعل مضارع مجزوم، وعلامة جزمه حذف حرف العلة (الياء) من آخره لأنه فعل معتل الآخر [4][8][9][14]، والفاعل ضمير مستتر وجوبًا تقديره «نحن».
مِن: حرف جر مبني على السكون، لا محل له من الإعراب. ذِكرى: اسم مجرور بـ«مِن»، وعلامة جره الكسرة المقدرة على الألف للتعذر، وهو مضاف. حَبيبٍ: مضاف إليه مجرور، وعلامة جره الكسرة الظاهرة. وَ: حرف عطف. مَنزِلِ: اسم معطوف على «حبيبٍ» مجرور مثله.
وقع الفعل المضارع في جواب الطلب؛ فالأمر «قِفا» تضمّن معنى الشرط، والتقدير: «إن تقفا نبكِ» [4][9][14].
حذف حرف العلة (الياء) من آخره، والكسرة الباقية على الكاف دليل على الياء المحذوفة [4][8]. وعلامات الجزم في العربية ثلاث: السكون، وحذف حرف العلة، وحذف النون [1][3][7][11][12].
This is the change log as documented by the corrected verdict itself:
The honest answer from this deliberation: the best is not a single model — it is the method. Qwen won the blind ranking yet still erred twice; Sonar gave the only textbook-perfect answer yet did not top the board; the lead model itself was wrong and got corrected. The fully correct answer came from no single model — it came from the system: five independent answers, a blind review, and a live-web fact-check with sources you can inspect.
If you are comparing AI models for Arabic, this is the comparison that actually matters: not impressions, but a deliberation documented with its date, results, and sources — repeatable any day, on any question.
Cited in the verdict
Numbers reflect the panel's retrieval order during fact-checking; only some sources were cited in the verdict. 15 sources were consulted in total, including video commentaries of the mu'allaqa not directly cited.