ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “NMT resources”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CLEmpiricalRecentJul 24, 2026

Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging

Abdullah Alabdullah, Arash Eslamighayour, Sarp Harbalioglu, Lifeng Han

This paper presents a study on improving healthcare-domain translation for four low-resource languages using Arabic and Persian as pivots, and introduces three transfer strategies.

View →
stat.MLcs.LGEmpiricalRecentJul 22, 2026

Non--negative matrix factorization using the \textit{R} package \textsf{nnmf}

Volkan Sevinç, Nikolas Kontemeniotis, Theodoros Perdikis, Michail Tsagris

This paper introduces a new R package for Non-negative Matrix Factorization (NMF) and compares its performance systematically with two other R packages using real-world data.

View →
cs.CLcs.AIcs.LGEmpiricalRecentJun 11, 2026

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

Marek Šuppa, Andrej Ridzik, Daniel Hládek, Natália Kňažeková +1 more

This paper introduces SkMTEB, a comprehensive text embedding benchmark for Slovak, and develops efficient, locally-deployable Slovak embeddings.

View →
cs.CLEmpiricalRecentJul 24, 2026

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

Varun Ghat Ravikumar, Sina Ahmadi, Lena Jäger, Rico Sennrich

This paper introduces a pipeline to extract grammatical rules, example sentences, and lexicons from grammar books and generates synthetic parallel corpora for fine-tuning machine translation models on…

View →
cs.CLRecentMay 29, 2026

Model-Based Quality Assessment for Massively Multilingual Parallel Data

Abdelaziz M. A. Ibrahim, Zihao Li, Jörg Tiedemann, Shaoxiong Ji

The paper proposes decomposing the assessment of massive multilingual parallel data into separate parallelism and quality estimation components, concluding that no single universal metric is reliable…

View →
cs.CLcs.AIRecentJun 1, 2026

KliniskVestBERT: BERT Model Specialised to Norwegian Clinical Texts

Christian Autenried, Cosimo Persia

This paper introduces KliniskVestBERT, a suite of BERT models specialized by pre-training on a large, diverse corpus of real-world Norwegian clinical texts, demonstrating superior performance for clin…

View →
cs.CLcs.AIcs.DBEmpiricalRecentJul 12, 2026

The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions

Filip Klubicka, Vasudevan Nedumpozhimana, Sneha Rautmare, Bora Caglayan +2 more

This paper explores the integration of several extensions for Natural Language to SQL (NL2SQL) translation, including NatSQL representation, preprocessing, fine-tuning, and reranker model.

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
cs.CLEmpiricalRecentJul 2, 2026

BamiBERT: A New BERT-based Language Model for Vietnamese

Dat Quoc Nguyen, Thinh Pham, Chi Tran, Linh The Nguyen

This paper introduces BamiBERT, a new Vietnamese language model based on BERT that addresses limitations of PhoBERT and sets a new state-of-the-art among base-sized Vietnamese encoders.

View →
cs.IRcs.CLDatasetRecentJun 23, 2026

PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation

Kirill Dubovikov, Omar El Mansouri, Hachem Madmoun, Yanda Li +11 more

This paper introduces PETRA, a large-scale Petroleum Engineering Text for Retrieval Adaptation dataset and pipeline that converts noisy public web data into a curated domain corpus and synthetic super…

View →
cs.AIcs.IRRecentMay 27, 2026

From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constraints

Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge

The paper introduces an LLM-based pipeline that tags learning resources with structured competencies, achieving strong performance while providing traceable evidence and leveraging graph constraints.

View →
cs.AIRecentMay 27, 2026

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

Xiang Wang, Tingting Zhang, Sen Wang, Ying Wu +3 more

The paper introduces PetroBench, a comprehensive benchmark for evaluating Large Language Models across various domains of petroleum engineering, finding that models perform better on subjective tasks…

View →
cs.LGcs.CVEmpiricalRecentJul 26, 2026

Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation

Anurag Roy, Riddhiman Moulick, Vinay Kumar Verma, Saptarshi Ghosh +1 more

This paper proposes a source-free continual test-time adaptation methodology using a controlled teacher adaptation and class prototype estimation.

View →
cs.IREmpiricalRecentJun 11, 2026

Hybrid Neural Retrieval with Generative Query Refinement for Quranic Passage Retrieval

Mohamed G. Salman, Mohammad E. Moftah, Ali Hamdi

This paper proposes a neural architecture for Quranic Passage Retrieval using hybrid candidate retrieval, semantic reranking, confidence gating, and multi-verse aggregation.

View →
cs.CLRecentMay 31, 2026

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

Xiaoqi He, Kaixin Lan, Mu You, Tao Fang +2 more

The paper proposes MACAT, a Multi-Agent Culture-Aware Translation framework, to selectively translate culture-loaded words in ancient Chinese texts, achieving superior performance over existing method…

View →
cs.CVcs.AIcs.MMEmpiricalRecentJul 10, 2026

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao +12 more

This paper presents the benefits of visual pretraining for foundation model intelligence, outperforming text-only pretraining on multiple backbones and benchmarks.

View →
cs.CLcs.AIcs.IREmpiricalRecentJun 23, 2026

MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

Junhyeok Lee, Han Jang, Hyeonjin Goh, Kyu Sung Choi

This paper introduces MMed-Bench-IR, a benchmark for multilingual medical retrieval in clinical settings, evaluating cross-lingual alignment, concept discrimination, and evidence retrieval.

View →
cs.CLcs.AIRecentMay 31, 2026

Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization

Sangwon Ryu, Yihong Liu, Mingyang Wang, Yunsu Kim +3 more

The paper introduces a new benchmark for multi-target cross-lingual summarization (MTXLS) and proposes an activation steering method that significantly improves LLM performance by guiding the generati…

View →
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View →