IEEE BigData 2025 · Workshop Archive

LLMs4All

LLMs, Big Data, and Multilinguality for All

A workshop on inclusive, multilingual and scalable language technologies, with particular attention to low-resource and underrepresented languages.

Workshop completed
10 December 202509:00–17:00, Macau local time
MacauIEEE BigData 2025
Hybrid workshopWhole-day programme
About LLMs4All

Inclusive language technology at the intersection of LLMs and Big Data

LLMs4All brought together research on large language models, multilinguality and scalable data methods, with a strong focus on languages and communities that remain underrepresented in mainstream NLP.

The workshop examined how large-scale data collection, cross-lingual transfer, efficient adaptation, retrieval-augmented generation and multilingual evaluation can improve the reach and reliability of language technologies in low-resource settings. It also provided a forum for practical work spanning text, speech, multimodal systems, domain adaptation and responsible AI.

Hosted as a whole-day hybrid workshop at IEEE BigData 2025 in Macau, LLMs4All is now maintained as an archive of the event, its programme, accepted papers and published proceedings.

21accepted workshop papers
1 dayfull workshop programme
HybridMacau and online participation
IEEEpublished conference proceedings

Proceedings are now published

The LLMs4All papers are available through the IEEE Computer Society Digital Library as part of the IEEE BigData 2025 proceedings.

IEEE proceedings
Workshop scope

Research themes

The programme covered methods, infrastructure, evaluation and applications for multilingual and low-resource language technologies.

Scalable data collection and curationLarge-scale multilingual datasets, annotation, web data and multimodal resources.
Cross-lingual and multilingual learningTransfer learning, multilingual embeddings, distillation and language-inclusive modelling.
Efficient and inclusive model trainingParameter-efficient fine-tuning, LoRA, PEFT, quantisation and compact models.
Retrieval-augmented generationGrounding multilingual LLMs in external structured and unstructured knowledge.
Multimodal language modelsSystems combining text, image, audio and other modalities across languages.
Big Data infrastructure and pipelinesDistributed processing, data engineering and scalable NLP infrastructure.
Responsible and culturally aware AIBias, fairness, representation, cultural context and equitable access.
Benchmarking and evaluationMetrics and benchmarks for robustness, faithfulness and multilingual generalisation.
Workshop contributions

Accepted papers

21 papers accepted for LLMs4All 2025
01

IDR-RAG: An Iterative Draft-Revision Agent-like RAG Pipeline for Efficient and Accurate Knowledge Retrieval in Large-Scale Private Domains

Xia Jiaju, Cao Lei

02

Polypersona: Persona-Grounded LLM for Synthetic Survey Responses

Tejaswani Dash, Dinesh Karri, Anudeep Vurity, Gautam Datla, Tazeem Ahmad, Saima Rafi, Rohith Tangudu

03

Evaluation of Large Language Models for Understanding Counterfactual Reasoning in Texts

S. I. M. Adnan, Abrar Hameem, Shikha Anirban, Md. Saiful Islam, Md. Musfique Anwar

04

How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation

Muskaan Chopra, Lorenz Sparrenberg, Sarthak Khanna, Rafet Sifa

05

Scaling Classical NLP Pipelines for Under-Resourced Old English: Character-Level Models, Unsupervised Pretraining, and Supervised Data Growth

Ana Elvira Ojanguren López, Javier Martín Arista, Darío Metola Rodríguez

06

Enhanced Old English NER via Morphology-Aware Analysis, Cross-Germanic Transfer, and Domain-Specific Patterns

Javier Martín Arista, Darío Metola Rodríguez, Daniel B. Morris

07

Benchmarking LLM Optimization Strategies for Clinical NER: A Comparative Analysis of DSPy GEPA against Domain-Specific Transformers

Justin Varghese, Yi Shang

08

Arabic Prompts with English Tools: A Benchmark

Konstantin Kubrak, Ahmed El-Moselhy, Ammar Alsulami, Remaz Altuwaim, Hassan Ismail Fawaz, Faisal Alsaby

09

AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs

Mo El-Haj, Paul Rayson

10

Temporal-Aware RAG for Multilingual ESG Document Retrieval: A Low-Resource Approach to Time-Sensitive Question Answering

Nguyen Anh Kiet Truong

11

UniFi-LLM: A Unified Large Language Model for Financial Data Generation and Fraud Prediction

Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli

12

Tackling Low-Resource K-12 Hand-Drawn Mathematics VQA: Unified Regularization with Compute-Aware Expert Token Architecture

Hai Li, Wanli Xing, Chenglu Li, Bailing Lyu

13

Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis

Felipe Ribeiro Fujita de Mello, Hideyuki Takada

14

Governance-Aware Hybrid Fine-Tuning for Multilingual Large Language Models

Haomin Qi, Chengbo Huang, Zihan Dai, Yunkai Gao

15

Leveraging LLM Agents for Autonomous Web Penetration Testing Targeting SQL Injection Vulnerability

Thanh Phong Tran, Le Bao Phuc Nguyen, Trong Nghia To, Van Hau Pham, The Duy Phan

16

XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs

Iñaki Lacunza, José Javier Saiz, Alexander Shvets, Aitor Gonzalez-Agirre, Marta Villegas

17

Arabic OCR in the Age of Multimodal Models: A Comprehensive Comparative Evaluation

Hossam Elsafty, Farizeh Aldabbas, Rafet Sifa

18

Copyright Infringement Issues and Mitigations in Data for Training Generative AI

Anna Arnaudo, Riccardo Coppola, Maurizio Morisio, Antonio Vetrò, Maurizio Borghi, Bryan Khan, Riccardo Raso

19

Cluster-aware Item Prompt Learning for Session-based Recommendation

Wooseong Yang, Chen Wang, Zihe Song, Weizhi Zhang, Philip S. Yu

20

Fine-tuning Large-Language-Models using Federated Learning & Blockchain

Soham Ratnaparkhi, Saeed Samet

21

Targeted Knowledge Enhancement: A Systematic Continual Pre-training Approach for Effective Domain Adaptation

Yiqun Wang, Chaoqun Wan, Xiang Tian, Xuesong Liu, Yaowu Chen

10 December 2025

Workshop programme

The final programme is retained here as an archive of the workshop day.

View the full workshop schedule
TimeTitlePresenter
09:00–09:15Evaluation of Large Language Models for Understanding Counterfactual Reasoning in TextsMd. Saiful Islam
09:15–09:30Scaling Classical NLP Pipelines for Under-Resourced Old English: Character-Level Models, Unsupervised Pretraining, and Supervised Data GrowthAna Elvira Ojanguren López
09:30–09:45Polypersona: Persona-Grounded LLM for Synthetic Survey ResponsesAnudeep Vurity
09:45–10:00IDR-RAG: An Iterative Draft-Revision Agent-like RAG PipelineCao Lei
10:00–10:15Benchmarking LLM Optimisation Strategies for Clinical NERJustin Varghese
10:15–10:30Targeted Knowledge EnhancementYiqun Wang
10:30–11:00Coffee break
11:00–11:15Enhanced Old English NERJavier Martín Arista
11:15–11:30Arabic Prompts with English Tools: A BenchmarkKonstantin Kubrak
11:30–11:45UniFi-LLMGiridhar Pamisetty
11:45–12:00Federated Learning & BlockchainSoham Ratnaparkhi
12:00–12:15LLM Agents for Web Penetration TestingThanh Phong Tran
12:15–12:30Temporal-Aware RAG for Multilingual ESGNguyen Anh Kiet Truong
12:30–14:00Lunch break
14:00–14:15AraFinNews: Arabic Financial SummarisationMo El-Haj
14:15–14:30XDoGE: Multilingual Data ReweightingIñaki Lacunza
14:30–14:45Copyright Infringement in Generative AIAnna Arnaudo
14:45–15:00Arabic OCR in the Age of Multimodal ModelsHossam Elsafty
15:00–15:15Compact LMs for On-Device MT Error DetectionMuskaan Chopra
15:15–15:30Better Data Selection for LLM Fine-TuningFelipe Ribeiro Fujita de Mello
15:30–16:00Coffee break
16:00–16:15Governance-Aware Hybrid Fine-TuningHaomin Qi
16:15–16:30Low-Resource K-12 Math VQAHai Li
16:30–16:45Cluster-aware Item Prompt LearningWooseong Yang
Organising committee

LLMs4All organisers

The workshop was organised by researchers from VinUniversity and the NLP @ VinUniversity Research Group.

LLMs4All workshop archive

For enquiries about the workshop or future NLP @ VinUniversity activities, contact the research group.

elhaj.m@vinuni.edu.vn  ·  vinnlp.com