LLMs, Big Data, and Multilinguality for All
A workshop on inclusive, multilingual and scalable language technologies, with particular attention to low-resource and underrepresented languages.
LLMs4All brought together research on large language models, multilinguality and scalable data methods, with a strong focus on languages and communities that remain underrepresented in mainstream NLP.
The workshop examined how large-scale data collection, cross-lingual transfer, efficient adaptation, retrieval-augmented generation and multilingual evaluation can improve the reach and reliability of language technologies in low-resource settings. It also provided a forum for practical work spanning text, speech, multimodal systems, domain adaptation and responsible AI.
Hosted as a whole-day hybrid workshop at IEEE BigData 2025 in Macau, LLMs4All is now maintained as an archive of the event, its programme, accepted papers and published proceedings.
The LLMs4All papers are available through the IEEE Computer Society Digital Library as part of the IEEE BigData 2025 proceedings.
The programme covered methods, infrastructure, evaluation and applications for multilingual and low-resource language technologies.
Xia Jiaju, Cao Lei
Tejaswani Dash, Dinesh Karri, Anudeep Vurity, Gautam Datla, Tazeem Ahmad, Saima Rafi, Rohith Tangudu
S. I. M. Adnan, Abrar Hameem, Shikha Anirban, Md. Saiful Islam, Md. Musfique Anwar
Muskaan Chopra, Lorenz Sparrenberg, Sarthak Khanna, Rafet Sifa
Ana Elvira Ojanguren López, Javier Martín Arista, Darío Metola Rodríguez
Javier Martín Arista, Darío Metola Rodríguez, Daniel B. Morris
Justin Varghese, Yi Shang
Konstantin Kubrak, Ahmed El-Moselhy, Ammar Alsulami, Remaz Altuwaim, Hassan Ismail Fawaz, Faisal Alsaby
Mo El-Haj, Paul Rayson
Nguyen Anh Kiet Truong
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli
Hai Li, Wanli Xing, Chenglu Li, Bailing Lyu
Felipe Ribeiro Fujita de Mello, Hideyuki Takada
Haomin Qi, Chengbo Huang, Zihan Dai, Yunkai Gao
Thanh Phong Tran, Le Bao Phuc Nguyen, Trong Nghia To, Van Hau Pham, The Duy Phan
Iñaki Lacunza, José Javier Saiz, Alexander Shvets, Aitor Gonzalez-Agirre, Marta Villegas
Hossam Elsafty, Farizeh Aldabbas, Rafet Sifa
Anna Arnaudo, Riccardo Coppola, Maurizio Morisio, Antonio Vetrò, Maurizio Borghi, Bryan Khan, Riccardo Raso
Wooseong Yang, Chen Wang, Zihe Song, Weizhi Zhang, Philip S. Yu
Soham Ratnaparkhi, Saeed Samet
Yiqun Wang, Chaoqun Wan, Xiang Tian, Xuesong Liu, Yaowu Chen
The final programme is retained here as an archive of the workshop day.
| Time | Title | Presenter |
|---|---|---|
| 09:00–09:15 | Evaluation of Large Language Models for Understanding Counterfactual Reasoning in Texts | Md. Saiful Islam |
| 09:15–09:30 | Scaling Classical NLP Pipelines for Under-Resourced Old English: Character-Level Models, Unsupervised Pretraining, and Supervised Data Growth | Ana Elvira Ojanguren López |
| 09:30–09:45 | Polypersona: Persona-Grounded LLM for Synthetic Survey Responses | Anudeep Vurity |
| 09:45–10:00 | IDR-RAG: An Iterative Draft-Revision Agent-like RAG Pipeline | Cao Lei |
| 10:00–10:15 | Benchmarking LLM Optimisation Strategies for Clinical NER | Justin Varghese |
| 10:15–10:30 | Targeted Knowledge Enhancement | Yiqun Wang |
| 10:30–11:00 | Coffee break | |
| 11:00–11:15 | Enhanced Old English NER | Javier Martín Arista |
| 11:15–11:30 | Arabic Prompts with English Tools: A Benchmark | Konstantin Kubrak |
| 11:30–11:45 | UniFi-LLM | Giridhar Pamisetty |
| 11:45–12:00 | Federated Learning & Blockchain | Soham Ratnaparkhi |
| 12:00–12:15 | LLM Agents for Web Penetration Testing | Thanh Phong Tran |
| 12:15–12:30 | Temporal-Aware RAG for Multilingual ESG | Nguyen Anh Kiet Truong |
| 12:30–14:00 | Lunch break | |
| 14:00–14:15 | AraFinNews: Arabic Financial Summarisation | Mo El-Haj |
| 14:15–14:30 | XDoGE: Multilingual Data Reweighting | Iñaki Lacunza |
| 14:30–14:45 | Copyright Infringement in Generative AI | Anna Arnaudo |
| 14:45–15:00 | Arabic OCR in the Age of Multimodal Models | Hossam Elsafty |
| 15:00–15:15 | Compact LMs for On-Device MT Error Detection | Muskaan Chopra |
| 15:15–15:30 | Better Data Selection for LLM Fine-Tuning | Felipe Ribeiro Fujita de Mello |
| 15:30–16:00 | Coffee break | |
| 16:00–16:15 | Governance-Aware Hybrid Fine-Tuning | Haomin Qi |
| 16:15–16:30 | Low-Resource K-12 Math VQA | Hai Li |
| 16:30–16:45 | Cluster-aware Item Prompt Learning | Wooseong Yang |
The workshop was organised by researchers from VinUniversity and the NLP @ VinUniversity Research Group.
General Chair
VinUniversity
Programme Chair
VinUniversity
Programme Chair
VinUniversity
Publication Chair
VinUniversity
Publication Chair
VinUniversity
Publicity Chair
VinUniversity
Publicity Chair
VinUniversity
Publicity Chair
VinUniversity
For enquiries about the workshop or future NLP @ VinUniversity activities, contact the research group.