Daily MT Picks

Archives
Subscribe
September 30, 2026

Machine Translation Digest for Sep 29 2026

A Greek papyrus HTR pipeline can still support four papyrological tasks even when character error rate is not near zero; in tests on 63,846 current edited readings, the paper maps where transcription quality starts to matter for document discovery and literary identification. Elsewhere, one study asks whether correcting a fixed set of poisoned rows is better than deleting them when fine-tuning Qwen2.5-14B-Instruct, another measures how GPT-generated edit-inducing questions affect revisions to ICLR and NeurIPS manuscripts, and MARCO frames molecular optimisation as a bounded proposal–feedback–revision problem for conditional editing.


Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision

Correct poisoned rows with the right answer rather than deleting them: replacing a quarter of bad medical examples cuts emergent misalignment by about a third, while deletion barely moves it.

Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter's. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.


Generating Edit-Inducing Questions for AI Research Manuscripts

For revision workflows, GPT-generated questions can drive larger edits than human reviewers, but only a small fraction are actually edit-inducing, so the prompt quality bar stays high.

We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.


MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

When a model must revise under feedback, train it on multi-turn proposal–feedback–revision trajectories: MARCO improves the property-similarity tradeoff on MuMOInstruct and still benefits from up to five turns at inference.

Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.


Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts

For Greek papyrus HTR, document-type classification tolerates about 7.5% CER, broad documentary-versus-literary labeling about 20%, and dating collapses near 3%, so task-specific error budgets matter.

Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, and how accurate they must be for a given papyrological task has not been examined. To answer this and set a benchmark for Greek papyrus HTR, we test a range of character error rates (CER) against four papyrological tasks, using published editions as ground truth. Methods: From 63,846 current editions of Greek texts in papyri.info, we imitate a letters-only "perfect HTR" output by removing the editorial layer, then degrade it with a seeded algorithm to exact CERs of 1 - 50%, with lost lines and four error-shape variants. On these data we train small models (TF-IDF, fastText, a character CNN, ByT5-small) for document type, dating and documentary-versus-literary classification, and apply eight keyword search methods. We compare models trained on clean text with models retrained at a specific CER level, and evaluate across CERs. Results: Tolerance differs by task. With clean-trained models, documentary-versus-literary classification retains 90% of its metric up to 20% CER; document type up to 7.5%; subtypes and search up to 5%; dating only up to 3%. Retraining on text containing character errors largely eliminates the sharp degradation that otherwise sets in above 15% CER. Models generally tolerate concentrated damage in a long document better than small errors spread across a short text. Conclusion: The study provides a CER target for each of the four tasks and shows that models trained on noisy text make current, imperfect text recognition useful for them.


Papers announced by arXiv on Tuesday, 29 September 2026, covering submissions from 28 September to 29 September 2026.

Curated by yukajii.com
Don't miss what's next. Subscribe to Daily MT Picks:
← Newer Machine Translation Digest for Sep 30 2026 Older → Machine Translation Digest for Sep 28 2026
Share this email:
Share on LinkedIn
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.