Diagnose hidden degradation
Identify whether apparent multilingual failure comes from comprehension, instruction adherence, output-language control or structured generation.
Instruction-Language Drift in Adapted Multilingual Instruction-Tuned LLMs
Fine-tuning a multilingual model to improve one task can silently change how it understands, follows and produces language elsewhere. PolyDrift studies where that drift appears, why it happens, and how it can be reduced.
PolyDrift treats multilingual capability as something that can change after post-adaptation. Instead of relying on a single performance score, the project diagnoses which part of instruction-following behaviour changes after fine-tuning.
A post-adaptation change in the model's ability to understand target-language input, follow instructions in the intended language, control the language of its output, or preserve required structured responses such as labels and JSON.
Identify whether apparent multilingual failure comes from comprehension, instruction adherence, output-language control or structured generation.
Examine whether adaptation gains in English come with reduced reliability in Vietnamese or Arabic, especially in real-world multilingual use.
Test lightweight stabilisation strategies that reduce drift while retaining the benefits of task adaptation.
The project moves from measuring drift, to locating the failure, explaining the internal change, and testing practical mitigation.
Quantify pre- and post-adaptation behaviour across English, Vietnamese and Arabic, with controlled changes to input, instruction and expected output language.
Separate failures in input understanding, instruction-language following, target-language output and structured response control rather than collapsing them into one score.
Connect behavioural drift with hidden-state movement, language-identity encoding and changes in cross-lingual alignment before and after adaptation.
Evaluate multilingual rehearsal, balanced multilingual instructions and representation anchoring as ways to reduce drift without sacrificing target-task gains.
A controlled workflow isolates behavioural and representation-level changes rather than treating multilingual performance as a static property.
Can the adapted model still understand target-language input when the instruction and expected label remain controlled?
Can it follow Vietnamese or Arabic instructions reliably after adaptation, independent of output-language demands?
Can it answer in the requested language, or does it drift back towards English or code-switch unexpectedly?
Can it preserve the required label, schema or JSON-style response format across multilingual prompt conditions?
English, Vietnamese and Arabic provide a meaningful test bed for studying adaptation across resource levels, scripts, morphology and multilingual representation.
The high-resource adaptation language and primary reference point for measuring gains and post-adaptation change.
Directly relevant to Vietnam and VinUniversity, and underrepresented in many multilingual LLM evaluation settings.
A typologically and orthographically distinct language with rich morphology, regional variation and broad multilingual NLP relevance.
PolyDrift is designed to leave behind reusable research infrastructure for analysing multilingual model adaptation beyond this project.
A controlled English–Vietnamese–Arabic instruction-language suite for testing input understanding, instruction following, output control and structured responses.
A baseline and post-adaptation evaluation pipeline with consistent prompt configurations, metrics, validation checks and experiment tracking.
Controlled model adapters or checkpoints where licences permit, together with systematic records of how behaviour changes across adaptation conditions.
Tools for task metrics, structured-output checks, language control analysis and representation-level comparison before and after adaptation.
A comparative assessment of multilingual rehearsal, balanced instructions and representation anchoring for reducing drift while preserving adaptation gains.
Public code where licences allow, benchmark and governance documentation, prompt templates, evaluation scripts and practical guidance for reliable multilingual adaptation.
The work plan keeps the core experiment focused while allowing selected extensions where data and computational resources permit.
Project setup, data governance, multilingual prompt construction, baseline evaluation and hidden-state activation caching.
Parameter-efficient adaptation, diagnostic evaluation across languages, structured output analysis and representation comparison.
Multilingual rehearsal, balanced instruction prompts and representation anchoring, followed by comparative drift-reduction analysis.
Software packaging, reproducible pipeline documentation, resource release where licences allow, research communication and future programme development.
PolyDrift brings together researchers in multilingual NLP, large language models, language resources, model adaptation and trustworthy AI.



The scientific contribution is a shift from asking whether multilingual performance drops to identifying precisely what changes after adaptation. This supports safer deployment in settings where multilingual systems are used for education, public communication, moderation and institutional decision support.
For research collaboration, technical discussion or information about the project, contact the PolyDrift Principal Investigator at VinUniversity.