Multilingual NLP · Vietnamese & English · Free-text analytics

FreeTxt-Vi

A benchmarked Vietnamese-English toolkit for segmentation, sentiment, summarisation and corpus exploration.

FreeTxt-Vi brings corpus linguistics and modern NLP together in an accessible web-based environment for building, exploring and interpreting bilingual free-text collections, without requiring users to write code.

Vietnamese + EnglishBilingual free-text analysis
Benchmarked NLPSegmentation, sentiment & summarisation
12-month projectLaunched May 2025
About FreeTxt-Vi

Making bilingual free-text analysis more accessible

Qualitative responses from surveys, questionnaires, interviews and feedback forms contain valuable information, but analysing large collections manually is slow and difficult to scale. FreeTxt-Vi provides an integrated environment for exploring Vietnamese-English text with corpus methods and NLP.

Văn Bản Tự Do

FreeTxt-Vi extends the Welsh-English FreeTxt research platform to Vietnamese-English data. The project develops a unified bilingual workflow for preparing text, analysing sentiment, generating summaries and exploring patterns in free-text collections.

The research places particular emphasis on Vietnamese, a widely spoken language that remains comparatively underrepresented in NLP resources and evaluation. The toolkit is designed for researchers and practitioners who work with qualitative text but may not have specialist programming or machine-learning expertise.

From raw responses to interpretable evidence

The aim is to make bilingual text analysis reproducible and practical by combining language-aware preprocessing, benchmarked NLP components, corpus exploration and interactive visualisation within a single workflow.

Designed for non-programmers

A web-based interface lowers the technical barrier for researchers who need to inspect and interpret qualitative text at scale.

Bilingual by design

Vietnamese and English are handled within one research workflow rather than as disconnected monolingual pipelines.

Evaluated components

The project benchmarks its core language-processing components rather than treating the toolkit as an unevaluated application layer.

Toolkit capabilities

A unified environment for analysing bilingual free text

FreeTxt-Vi combines established corpus-linguistic methods with transformer-based NLP so users can move between close inspection of text and scalable automatic analysis.

Vietnamese-English segmentation

A language-aware segmentation pipeline prepares Vietnamese and English text for downstream analysis, including a hybrid VnCoreNLP and Byte Pair Encoding strategy for Vietnamese.

VnCoreNLPBPEBilingual preprocessing

Sentiment analysis

Transformer-based sentiment modelling supports scalable interpretation of opinions and attitudes expressed across Vietnamese-English free-text collections.

Transformer modelsFine-tuningEvaluation

Abstractive summarisation

LLM-based summarisation helps users condense larger bodies of responses into shorter, interpretable summaries while retaining the key content of the source collection.

Qwen2.5LLMsBilingual summaries

Corpus exploration & visualisation

Concordancing, keyword analysis, word-relation exploration and interactive visualisation support deeper inspection of recurring themes, language patterns and relationships.

ConcordanceKeywordsVisual analytics
Research design

From bilingual text collection to interpretable analysis

The project treats FreeTxt-Vi as a set of language-processing components that can be individually evaluated, integrated and used within a broader corpus-analysis workflow.

CollectImport bilingual free-text responses
SegmentPrepare Vietnamese-English text
AnalyseSentiment and corpus-level patterns
SummariseGenerate concise abstractive summaries

Segmentation benchmark

Tests the reliability of the bilingual preprocessing pipeline, including Vietnamese word segmentation and tokenisation choices.

Sentiment benchmark

Measures classification quality against established baselines to assess whether the toolkit provides robust bilingual sentiment analysis.

Summarisation benchmark

Evaluates the LLM-based summarisation component in Vietnamese and English rather than relying solely on qualitative demonstration.

Project outputs

Research, software and reusable methodology

FreeTxt-Vi produces both an applied bilingual analysis environment and research evidence on the language-processing components that underpin it.

LREC 2026Project publication

FreeTxt-Vi: A Benchmarked Vietnamese-English Toolkit for Segmentation, Sentiment, and Summarisation

Nguyen Huy, H., El-Haj, M., Knight, D., and Rayson, P.

Language Resources and Evaluation Conference (LREC), 2026. The paper presents the FreeTxt-Vi architecture and a three-part evaluation of segmentation, sentiment analysis and summarisation across Vietnamese and English.

Open-source toolkit

A web-based environment for creating and analysing Vietnamese-English text collections through corpus and NLP methods.

Benchmark methodology

A reproducible evaluation setup for testing segmentation, sentiment and summarisation as distinct components of a bilingual NLP toolkit.

Vietnamese language resources

The project contributes practical methods, evaluation evidence and reusable workflows for a language that remains underrepresented in mainstream NLP.

Research team

FreeTxt-Vi team

The project connects multilingual NLP, corpus linguistics and language-resource expertise across VinUniversity, Lancaster University and Cardiff University.

Mo El-Haj

Mo El-Haj

Principal Investigator

VinUniversity, Vietnam

Multilingual NLP, low-resource language technology and LLM-based text analysis.

Profile
Paul Rayson

Paul Rayson

Co-Investigator

Lancaster University, UK

Corpus linguistics, NLP evaluation and language-resource development.

Profile
Dawn Knight

Dawn Knight

Co-Investigator

Cardiff University, UK

Corpus linguistics, digital language resources and the original FreeTxt research programme.

Profile
Nguyen Huy Hung

Nguyen Huy Hung

Research Assistant

VinUniversity, Vietnam

Vietnamese-English NLP pipeline development, experimentation and benchmarking.

LinkedIn
Collaboration

Institutions and research groups

FreeTxt-Vi develops Vietnamese-English language technology through collaboration across VinUniversity, UCREL at Lancaster University and Cardiff University.

Why it matters

Supporting research with multilingual qualitative data

FreeTxt-Vi is designed for settings where rich qualitative text is common but difficult to process at scale. By reducing technical barriers, it supports reproducible bilingual analysis and the development of Vietnamese NLP resources.

Education & social researchAnalyse open-ended survey responses, feedback and qualitative research data across Vietnamese and English.
Digital humanities & cultural heritageExplore bilingual collections using corpus methods, visualisation and language-aware NLP components.
Inclusive language technologyStrengthen tools and evaluation practices for Vietnamese within multilingual NLP research.

Get in touch

Project lead: Dr Mo El-Haj

Address: CECS, VinUniversity, Hanoi, Vietnam

Email: elhaj.m@vinuni.edu.vn

Research group: NLP @ VinUniversity

Send a message