Daily MT Picks

Archives
Subscribe
September 29, 2026

Machine Translation Digest for Sep 28 2026

How do you score a natural-language-to-logic translation when a single quantifier flip can break every proof downstream? One paper answers with SIV, a prover-grounded metric for NL→FOL that probes candidate formulas against theorem-checked targets instead of leaning on BLEU, BERTScore, or Smatch++; another introduces BaatCheet, a multilingual corpus of roughly 49,000 Indian-language dialogues for dialogue translation, where informality, speaker interaction, and discourse coherence matter in ways sentence-level benchmarks miss.


Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL$\rightarrow$FOL

Use prover-backed probes like SIV when evaluating NL-to-FOL translation, because it separates dropped content from over-assertion and tracks error severity far better than BLEU or Smatch++.

A standard pipeline for symbolic reasoning over natural-language problems translates them into first-order logic and invokes a theorem prover. The translation step is the bottleneck: swap "every" for "some" and every inference that follows is corrupted. Yet today's metrics often score more broken translations higher than less broken ones, because BLEU, BERTScore, and Smatch++ reward surface overlap that the worst errors happen to preserve. We introduce SIV, which derives two kinds of probes from the target formula and uses a theorem prover to verify the candidate translation against each. Positive probes are statements the candidate must entail, which detect translations that drop content; contrastive probes are statements the candidate must not entail, which detect translations that assert more than the original. On a controlled pool of perturbed FOLIO translations, the severity of the error accounts for 80% of SIV's score variance, compared with at most 17% for any prior metric. Across six error classes on a disjoint pool, SIV scores the reference above the perturbed candidate in over 99% of pairs. Because each probe is labeled with what it tests, the failure pattern also supplies a labeled error trace, recovering the perturbation class at macro-F1 0.638, nearly double the score-only baseline. On 434 expert-audited real LLM translations, SIV attains the top AUC, uniquely detects and grades expert-labeled major errors, and abstains, rather than mis-scoring, on out-of-vocabulary translations.


BaatCheet: A Multilingual Corpus for Dialogue Translation in Indian Languages

BaatCheet gives you about 49,000 multilingual dialogue examples for Indian-language translation, and fine-tuning open-source LLMs beats zero- and few-shot baselines on this conversational setting.

Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks focus on sentence-level or formal text, making it difficult to assess translation quality of the dialogue phenomena. In this work, we introduce BaatCheet, a multilingual dialogue corpus named after the Hindi term for conversation or chitchat, containing approximately 49,000 dialogues for dialogue translation across five translation directions. We fine-tune five open-source LLMs across seven training data configurations and find that fine-tuning yields substantial gains over zero- and few-shot baselines. To comprehensively evaluate dialogue translation quality, we employ multiple evaluation strategies, including automatic metrics, LLM-as-judge, and human assessments using an SQM-guided Direct Assessment (DA) Protocol.


NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech

NVAlign matters if you localise TTS output with inline non-verbal tags, because direct-gradient post-training improves tag-following without sacrificing speaker similarity or speech quality.

While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an NVV-aware automatic speech recognition (NV-ASR) model on NVV-annotated speech, then freeze the NV-ASR model to serve as the reward model for post-training. A two-step gradient surrogate enables efficient reward backpropagation through the flow-matching sampler to jointly update the autoregressive backbone and acoustic flow head. Fidelity penalties and reference-velocity regularization help preserve speaker similarity and speech quality. Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines. These findings demonstrate that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS. Audio samples are available at https://nvalign.github.io/.


MixDetect: Word-Level Localization and Quantification of AI Editing

MixDetect is useful when you need to audit edited text, because it localises AI-written words and estimates edit intensity separately instead of collapsing both into one authorship label.

Large language models are increasingly used to edit human-written text rather than generate entire texts from scratch. Conventional AI-text detectors mainly distinguish human-written from fully AI-generated text, while recent methods for AI-edited text typically provide only a text-level label or editing-degree score. We introduce MixDetect, a word-level framework for localizing and quantifying AI editing. MixDetect separately predicts whether each word has been edited and, conditional on editing, how substantial the edit is, allowing editing scope and editing intensity to be estimated separately. During training, source--edited pairs are aligned to construct word-level supervision, while inference requires only the input text. Experiments show that MixDetect accurately localizes AI-edited words, reflects differences in editing intensity, and reveals different scope--intensity patterns across editing degrees and operations. The overall AI editing magnitude increases under additional AI editing, decreases when AI-generated text is edited by humans, and remains nearly unchanged under ordinary human-to-human editing. The aggregated text-level predictions also perform well on binary and ternary AI-text classification and remain effective under domain and generator shifts. These results show that AI editing can be analyzed beyond a single authorship label or editing-degree score by identifying both where AI editing occurs and how substantial the edits are.


Program-Verified Self-Evolution for Vision-Language Models

VQS replaces noisy majority-vote labelling with program-verified answer generation, so self-training uses claim-level checks and cuts label error from 24% to 6% in human review.

Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24% of majority-vote labels and 18% of model-judge labels produced during self-evolution are wrong. To address this problem, we present Verifiable QA Generation for Self-Evolving Models (VQS), which changes how the model judges answers. Instead of voting on an answer, the model parses each image into a structured record, such as a scene graph, a chart table, or a diagram graph. Fixed programs then write a question from the record and compute its answer. The model still acts as a visual checker, but it only confirms the individual facts the program reads, one short claim at a time. These claim-level checks select the parser's training targets, so the parser also improves without labels. Human raters find 94% of VQS answers correct, against 76% for majority voting. Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales and outperforms the strongest self-evolving baseline at each. Gains keep growing over three training rounds, reaching 3.84 points at 2B. Code is released at https://github.com/ahmedheakl/VQS


Papers announced by arXiv on Monday, 28 September 2026, covering submissions from 25 September to 28 September 2026.

Curated by yukajii.com
Don't miss what's next. Subscribe to Daily MT Picks:
← Newer Machine Translation Digest for Sep 29 2026 Older → Machine Translation Digest for Sep 27 2026
Share this email:
Share on LinkedIn
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.