Publications
The publications are grouped into two sections: conference and journal papers.
* Denotes Equal Contribution
- Findings EACL 2026MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level AssessmentOmid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, and Mahdieh Soleymani BaghshahIn Findings of the Association for Computational Linguistics: EACL 2026, 2026
Recent advancements in large vision-language models (VLMs) have primarily focused on English, with limited attention given to other languages. To address this gap, we introduce MEENA (also known as PersianMMMU), the first dataset designed to evaluate Persian VLMs across scientific, reasoning, and human-level understanding tasks. Our dataset comprises approximately 7,500 Persian and 3,000 English questions, covering a wide range of topics such as reasoning, mathematics, physics, diagrams, charts, and Persian art and literature. Key features of MEENA include: (1) diverse subject coverage spanning various educational levels, from primary to upper secondary school, (2) rich metadata, including difficulty levels and descriptive answers, (3) original Persian data that preserves cultural nuances, (4) a bilingual structure to assess cross-linguistic performance, and (5) a series of diverse experiments assessing various capabilities, including overall performance, the model’s ability to attend to images, and its tendency to generate hallucinations. We hope this benchmark contributes to enhancing VLM capabilities beyond English.
@inproceedings{ghahroodi2026meena, title = {{MEENA} ({PersianMMMU}): Multimodal-Multilingual Educational Exams for N-level Assessment}, author = {Ghahroodi, Omid and Hemmat, Arshia and Nouri, Marzia and Hosseini, Seyed Mohammad Hadi and Dastgheib, Doratossadat and Sanian, Mohammad Vali and Sahebi, Alireza and Zohrabi, Reihaneh and Rohban, Mohammad Hossein and Asgari, Ehsaneddin and Baghshah, Mahdieh Soleymani}, booktitle = {Findings of the Association for Computational Linguistics: EACL 2026}, year = {2026}, url = {https://aclanthology.org/2026.findings-eacl.340/}, doi = {10.18653/v1/2026.findings-eacl.340}, keyword = {conference} } - Findings ACL 2026Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language ModelsMohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, and Ehsaneddin AsgariIn Findings of the Association for Computational Linguistics: ACL 2026, 2026
Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence. Most existing evaluations focus on piecemeal or disconnected tasks, obscuring critical cognitive weaknesses and providing little insight for targeted improvement. To address this gap, we introduce BloomBench, part of the Almieyar benchmarking series, the first cognitively human-grounded, bilingual (English–Arabic) multimodal benchmark for VLMs. Grounded in Bloom’s Taxonomy, BloomBench systematically evaluates six levels of cognition (Remember, Understand, Apply, Analyze, Evaluate, Create) through carefully designed image–question–answer tasks. Built with a semi-automated pipeline and validated through a stratified hybrid quality assurance protocol, it ensures scalability, cultural inclusivity, and linguistic fidelity. Leveraging this framework, we conduct a comprehensive study of state-of-the-art VLMs to diagnose their cognitive profiles. Our analysis reveals a sharp cognitive asymmetry: while state-of-the-art models achieve strong performance ceilings in semantic understanding, they struggle substantially with factual recall and creative synthesis. This demonstrates that current general multimodal proficiency masks deeper limitations in specific cognitive layers. Furthermore, our study highlights a critical performance gap between Arabic and English, exposing limitations in current cross-lingual multimodal reasoning. These findings establish a foundation for developing more cognitively aligned and inclusive VLMs. The benchmark framework and dataset is available at: https://github.com/qcri/Almieyar-Oryx-BloomBench.
@inproceedings{abootorabi2026almieyar, title = {Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models}, author = {Abootorabi, Mohammad Mahdi and Ghahroodi, Omid and Madkoor, Anas and Nouri, Marzia and Dastgheib, Doratossadat and Asgari, Ehsaneddin}, booktitle = {Findings of the Association for Computational Linguistics: ACL 2026}, year = {2026}, url = {https://aclanthology.org/2026.findings-acl.1416/}, doi = {10.18653/v1/2026.findings-acl.1416}, keyword = {conference} } - Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri, and Hassan SajjadIn Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024), 2024
Interpreting and understanding the predictions made by deep learning models poses a formidable challenge due to their inherently opaque nature. Many previous efforts aimed at explaining these predictions rely on input features, specifically, the words within NLP models. However, such explanations are often less informative due to the discrete nature of these words and their lack of contextual verbosity. To address this limitation, we introduce the Latent Concept Attribution method (LACOAT), which generates explanations for predictions based on latent concepts. Our foundational intuition is that a word can exhibit multiple facets, contingent upon the context in which it is used. Therefore, given a word in context, the latent space derived from our training process reflects a specific facet of that word. LACOAT functions by mapping the representations of salient input words into the training latent space, allowing it to provide latent context-based explanations of the prediction.
@inproceedings{yu2024latentconceptbasedexplanationnlp, title = {Latent Concept-based Explanation of NLP Models}, author = {Yu, Xuemin and Dalvi, Fahim and Durrani, Nadir and Nouri, Marzia and Sajjad, Hassan}, booktitle = {Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024)}, year = {2024}, url = {https://arxiv.org/abs/2404.12545}, archiveprefix = {arXiv}, eprint = {2404.12545}, primaryclass = {cs.CL}, keyword = {conference} } - Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?Omid Ghahroodi, Marzia Nouri*, Mohammad Vali Sanian*, Alireza Sahebi*, Doratossadat Dastgheib, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah, and Mohammad Hossein RohbanIn First Conference on Language Modeling, 2024
Evaluating Large Language Models (LLMs) is challenging due to their generative nature, necessitating precise evaluation methodologies. Additionally, non-English LLM evaluation lags behind English, resulting in the absence or weakness of LLMs for many languages. In response to this necessity, we introduce Khayyam Challenge (also known as PersianMMLU), a meticulously curated collection comprising 20,805 four-choice questions sourced from 38 diverse tasks extracted from Persian examinations, spanning a wide spectrum of subjects, complexities, and ages. The primary objective of the Khayyam Challenge is to facilitate the rigorous evaluation of LLMs that support the Persian language. Distinctive features of the Khayyam Challenge are (i) its comprehensive coverage of various topics, including literary comprehension, mathematics, sciences, logic, intelligence testing, etc aimed at assessing different facets of LLMs such as language comprehension, reasoning, and information retrieval across various educational stages, from lower primary school to upper secondary school (ii) its inclusion of rich metadata such as human response rates, difficulty levels, and descriptive answers (iii) its utilization of new data to avoid data contamination issues prevalent in existing frameworks (iv) its use of original, non-translated data tailored for Persian speakers, ensuring the framework is free from translation challenges and errors while encompassing cultural nuances (v) its inherent scalability for future data updates and evaluations without requiring special human effort. Previous works lacked an evaluation framework that combined all of these features into a single comprehensive benchmark. Furthermore, we evaluate a wide range of existing LLMs that support the Persian language, with statistical analyses and interpretations of their outputs. We believe that the Khayyam Challenge will improve advancements in LLMs for the Persian language by highlighting the existing limitations of current models, while also enhancing the precision and depth of evaluations on LLMs, even within the English language context.
@inproceedings{ghahroodi2024khayyam, title = {Khayyam Challenge (Persian{MMLU}): Is Your {LLM} Truly Wise to The Persian Language?}, author = {Ghahroodi, Omid and Nouri, Marzia and Sanian, Mohammad Vali and Sahebi, Alireza and Dastgheib, Doratossadat and Asgari, Ehsaneddin and Baghshah, Mahdieh Soleymani and Rohban, Mohammad Hossein}, booktitle = {First Conference on Language Modeling}, year = {2024}, url = {https://openreview.net/forum?id=yIEyHP7AvH}, keyword = {conference} } - The Language Model, Resources, and Computational Pipelines for the Under-Resourced Iranian AzerbaijaniMarzia Nouri*, Mahsa Amani*, Reihaneh Zohrabi, and Ehsaneddin AsgariIn Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), Nov 2023
Iranian Azerbaijani is a dialect of the Azerbaijani language spoken by more than 16% of the population in Iran (>14 million). Unfortunately, a lack of computational resources is one of the factors that puts this language and its rich culture at risk of extinction. This work aims to create fundamental natural language processing (NLP) resources and pipelines for the processing and analysis of Iranian Azerbaijani introducing standard datasets and starter models for various NLP tasks such as language modeling, text classification, part-of-speech (POS) tagging, and machine translation. The proposed resources have been curated and preprocessed to facilitate the development of NLP models for Iranian Azerbaijani and provide a strong baseline for further research and development. This study is an example of bridging the gap in NLP for low-resource languages and promoting the advancement of language technologies in underrepresented languages. To the best of our knowledge, for the first time, this paper presents major infrastructures for the processing and analysis of Iranian Azerbaijani, with the ultimate goal of improving communication and information access for millions of individuals. Furthermore, our translation model’s online demo is accessible at https://azeri.parsi.ai/.
@inproceedings{nouri-etal-2023-language, title = {The Language Model, Resources, and Computational Pipelines for the Under-Resourced {I}ranian {A}zerbaijani}, author = {Nouri, Marzia and Amani, Mahsa and Zohrabi, Reihaneh and Asgari, Ehsaneddin}, editor = {Park, Jong C. and Arase, Yuki and Hu, Baotian and Lu, Wei and Wijaya, Derry and Purwarianti, Ayu and Krisnadhi, Adila Alfa}, booktitle = {Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 2: Short Papers)}, month = nov, year = {2023}, address = {Nusa Dua, Bali}, publisher = {Association for Computational Linguistics}, url = {https://aclanthology.org/2023.ijcnlp-short.19}, doi = {10.18653/v1/2023.ijcnlp-short.19}, pages = {166--174}, keyword = {conference} }
- Agents in the Wild 2026WebArena-Pro: A Heterogeneous, Multimodal, Reproducible Benchmark for Web AgentsImene Kerboua, Fatemeh Pesaran Zadeh, Xing Han Lù, Weijian Qi, Alexander Miller, Junyi Song, Yunjia Tian, Dongjin Kang, Seyeon Choi, Marzia Nouri, Ewen Gueguen, Matteo Boglioni, Fengyuan Liu, Zeyi Liao, Mengqi Yuan, Yue Li, Alexandre Lacoste, Alexandre Drouin, Spandana Gella, Huan Sun, Gunhee Kim, and Siva ReddyIn Second Workshop on Agents in the Wild: Safety, Security, and Beyond, 2026
Web agents powered by large language and vision-language models are increasingly applied to realistic browser work that spans heterogeneous applications, multimodal content, and stateful workflows. However, existing reproducible web-agent benchmarks cover only a small number of web applications drawn from a few software categories, and restrict modality to text and vision. Live benchmarks broaden site coverage but sacrifice reproducibility, since pages and data drift between runs. Moreover, existing benchmarks do not meaningfully evaluate whether agents can understand and use audio and video content embedded within web tasks. To address these gaps, we introduce WebArena-Pro, a benchmark comprising 300 tasks across 20 self-hosted web applications in six domain categories, spanning distinct interface conventions, workflows, and data models. Across the evaluated agents, the best performance is achieved by Gemini 3.1 Pro, which attains 37.0% success under a 50-step budget, while open-source models’ performance does not exceed 27.7% success. Among reproducible, human-curated web agent benchmarks, WebArena-Pro provides the broadest application coverage and the most comprehensive multimodal support to date. The benchmark treats audio and video as core observations alongside text and vision, with dedicated actions for extracting information from each. WebArena-Pro runs each task in isolation and supports reproducible, parallel evaluation. Tasks are authored through a dedicated annotator interface, filtered by LLM-assisted triage, and finally validated by humans before release.
@inproceedings{kerboua2026webarenapro, title = {WebArena-Pro: A Heterogeneous, Multimodal, Reproducible Benchmark for Web Agents}, author = {Kerboua, Imene and Zadeh, Fatemeh Pesaran and L{\`u}, Xing Han and Qi, Weijian and Miller, Alexander and Song, Junyi and Tian, Yunjia and Kang, Dongjin and Choi, Seyeon and Nouri, Marzia and Gueguen, Ewen and Boglioni, Matteo and Liu, Fengyuan and Liao, Zeyi and Yuan, Mengqi and Li, Yue and Lacoste, Alexandre and Drouin, Alexandre and Gella, Spandana and Sun, Huan and Kim, Gunhee and Reddy, Siva}, booktitle = {Second Workshop on Agents in the Wild: Safety, Security, and Beyond}, year = {2026}, url = {https://openreview.net/forum?id=eMuJZXwAn1}, keyword = {workshop} }
- Linguistic Resources and Transformer-based Models for the Machine Translations between Luri and Yazdi Dialects versus Standard PersianZahra Bahmani, Mohaddeseh Mirbeygi, Negin Hashemi Dijujin, Marzia Nouri, Mahsa Amani, Ehsan Asgari, Mahdieh Soleymani Baghshah, Hamid Beigy, Ali Movaghar, and Afzal MoghimiLanguage and Linguistics, 2022
Despite recent advances in developing language technologies for the standard Persian dialect, the official Iranian language, a large number of Iranian language variations remained computationally unexplored. Iranian languages, e.g., Kurdi, Azeri, and many Persian dialects are examples of low-resource language distinctions lacking significant linguistic resources such as machine-readable lexicons or part-of-speech (POS) taggers. Efforts in developing language technologies for such languages can significantly contribute to language survival in the digital era and promote cultural diversity. To the best of our knowledge, for the first time, we created linguistic resources for the Luri and the Yazdi dialects by introducing the first parallel corpora between these language variations and the modern Persian language. In this study, we train neural encoder-decoders (1) recurrent sequence-to-sequence and (2) transformer-based machine translation models and evaluate the trained model using BLEU score on an unseen test dataset.Availability of datasets and models: Datasets are available here at https://github.com/language-ml/dataset_yazdi_luri.git