Daily MT Picks

Archives
Subscribe
October 2, 2026

Machine Translation Digest for Oct 01 2026

One paper argues that a rule-based system still beats fine-tuned neural models for Polish-Silesian, while another takes the opposite tack and tries to improve LLMs at the tokenization level with Counting and Filtering plus Min-Cost Encoding. The dialect work tests on SiLTT, BOUQuET and FLORES, whereas the tokenization paper focuses on reducing sequence length and inference cost for large language models by globally minimizing segmentation cost.


Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models

For Polish-Silesian work, a rule-based system beats both open and commercial neural models on the new SiLTT testset, so dialect translation still needs curated data and explicit linguistic coverage.

Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, evaluating on SiLTT - a new Pol-Szl testset, alongside established BOUQuET and FLORES benchmarks. Results show our rule-based system is consistently strongest on SiLTT and BOUQuET datasets and that TranslateGemma fine-tuned on a curated dataset improves over strong neural baselines but does not surpass the rule-based system in dialectal settings. We release SiLTT and our best neural model to support further research.


Counting and Min-Cost Encoding for Tokenization in Large Language Models

Counting and Min-Cost Encoding cuts token counts versus BPE, with 26% to 30% better compression at 250K vocabulary and over 60% better token efficiency at 1M, which directly lowers inference cost.

Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training approach named Counting and Filtering (CNF) and a text encoding algorithm called Min-Cost Encoding (MCE). MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost. CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE. The CNF-MCE conbination offers several advantages over BPE, including higher token efficiency, greater scalability, and lower dependency. Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers. With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers. Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%. The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others. Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks. These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.


Papers announced by arXiv on Thursday, 1 October 2026, covering submissions from 30 September to 1 October 2026.

Curated by yukajii.com
Don't miss what's next. Subscribe to Daily MT Picks:
Older → Machine Translation Digest for Sep 30 2026
Share this email:
Share on LinkedIn
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.