Designed for non-programmers
A web-based interface lowers the technical barrier for researchers who need to inspect and interpret qualitative text at scale.
A benchmarked Vietnamese-English toolkit for segmentation, sentiment, summarisation and corpus exploration.
FreeTxt-Vi brings corpus linguistics and modern NLP together in an accessible web-based environment for building, exploring and interpreting bilingual free-text collections, without requiring users to write code.
Qualitative responses from surveys, questionnaires, interviews and feedback forms contain valuable information, but analysing large collections manually is slow and difficult to scale. FreeTxt-Vi provides an integrated environment for exploring Vietnamese-English text with corpus methods and NLP.
FreeTxt-Vi extends the Welsh-English FreeTxt research platform to Vietnamese-English data. The project develops a unified bilingual workflow for preparing text, analysing sentiment, generating summaries and exploring patterns in free-text collections.
The research places particular emphasis on Vietnamese, a widely spoken language that remains comparatively underrepresented in NLP resources and evaluation. The toolkit is designed for researchers and practitioners who work with qualitative text but may not have specialist programming or machine-learning expertise.
The aim is to make bilingual text analysis reproducible and practical by combining language-aware preprocessing, benchmarked NLP components, corpus exploration and interactive visualisation within a single workflow.
A web-based interface lowers the technical barrier for researchers who need to inspect and interpret qualitative text at scale.
Vietnamese and English are handled within one research workflow rather than as disconnected monolingual pipelines.
The project benchmarks its core language-processing components rather than treating the toolkit as an unevaluated application layer.
FreeTxt-Vi combines established corpus-linguistic methods with transformer-based NLP so users can move between close inspection of text and scalable automatic analysis.
A language-aware segmentation pipeline prepares Vietnamese and English text for downstream analysis, including a hybrid VnCoreNLP and Byte Pair Encoding strategy for Vietnamese.
Transformer-based sentiment modelling supports scalable interpretation of opinions and attitudes expressed across Vietnamese-English free-text collections.
LLM-based summarisation helps users condense larger bodies of responses into shorter, interpretable summaries while retaining the key content of the source collection.
Concordancing, keyword analysis, word-relation exploration and interactive visualisation support deeper inspection of recurring themes, language patterns and relationships.
The project treats FreeTxt-Vi as a set of language-processing components that can be individually evaluated, integrated and used within a broader corpus-analysis workflow.
Tests the reliability of the bilingual preprocessing pipeline, including Vietnamese word segmentation and tokenisation choices.
Measures classification quality against established baselines to assess whether the toolkit provides robust bilingual sentiment analysis.
Evaluates the LLM-based summarisation component in Vietnamese and English rather than relying solely on qualitative demonstration.
FreeTxt-Vi produces both an applied bilingual analysis environment and research evidence on the language-processing components that underpin it.
Language Resources and Evaluation Conference (LREC), 2026. The paper presents the FreeTxt-Vi architecture and a three-part evaluation of segmentation, sentiment analysis and summarisation across Vietnamese and English.
A web-based environment for creating and analysing Vietnamese-English text collections through corpus and NLP methods.
A reproducible evaluation setup for testing segmentation, sentiment and summarisation as distinct components of a bilingual NLP toolkit.
The project contributes practical methods, evaluation evidence and reusable workflows for a language that remains underrepresented in mainstream NLP.
The project connects multilingual NLP, corpus linguistics and language-resource expertise across VinUniversity, Lancaster University and Cardiff University.
VinUniversity, Vietnam
Multilingual NLP, low-resource language technology and LLM-based text analysis.
Profile
Lancaster University, UK
Corpus linguistics, NLP evaluation and language-resource development.
Profile
Cardiff University, UK
Corpus linguistics, digital language resources and the original FreeTxt research programme.
Profile
VinUniversity, Vietnam
Vietnamese-English NLP pipeline development, experimentation and benchmarking.
LinkedInFreeTxt-Vi develops Vietnamese-English language technology through collaboration across VinUniversity, UCREL at Lancaster University and Cardiff University.
FreeTxt-Vi is designed for settings where rich qualitative text is common but difficult to process at scale. By reducing technical barriers, it supports reproducible bilingual analysis and the development of Vietnamese NLP resources.
Project lead: Dr Mo El-Haj
Address: CECS, VinUniversity, Hanoi, Vietnam
Email: elhaj.m@vinuni.edu.vn
Research group: NLP @ VinUniversity