Horizon Lens — 30 September 2026
GPT-6.1 Sol makes a lower-cost pitch for agent work
OpenAI introduced GPT-6.1 Sol on 29 September, positioning it closer to Astra for coding, computer use and professional work. Standard API pricing is $2 per million input tokens and $10 per million output tokens; cached input is $0.10 per million. Those are API token prices, not a statement about subscription usage allowances or the cost of every completed task.
OpenAI says Sol 6.1 improves on its predecessor across its evaluations. On one factuality test, the low-effort error rate fell from 11.4% to 7.7%. The company explicitly says the prompts were selected from conversations where users had flagged errors and are not representative of ordinary use. These are vendor-reported results, not universal accuracy figures.
Analysis
Lower token prices can make a capable agent more affordable, but retries, context size and human corrections still matter. A cheaper response is not automatically a cheaper successful workflow.
Action
Compare models on the same bounded task and score the finished result. Record total usage and review time before changing a dependable workflow solely because a new model is available. Keep the old result for comparison, so an apparent saving is not hiding extra errors or unfinished steps.
NVIDIA brings in-context prediction to tables
NVIDIA’s new Kumo Tabular model predicts categories or numerical values from a table of labelled examples, without task-specific weight training. Its 29 September announcement describes three sizes, from 28 million to 215 million parameters, pretrained entirely on artificial tables. Code and weights are available, with commercial use under the OpenMDW-1.1 licence.
NVIDIA reports leading results on four benchmarks, but its own limitations section deserves equal attention. The model works on numerical and categorical columns; other inputs need conversion to features. Accuracy may degrade outside the training ranges or when the rows being predicted differ in distribution from the examples. The company recommends checking accuracy and calibration on held-out data.
Analysis
The potential saving is in preparing a new prediction task, not removing responsibility for testing it. Faster setup makes careful comparison easier; it does not make a benchmark representative of your records.
Action
Start with a historical task whose answers are already known. Keep evaluation rows separate and compare against your existing method before letting predictions influence real decisions.
A true fact can still carry the wrong citation
Multiverse Computing’s 29 September article introduces ProvenanceGuard, a verification layer for agents using multiple tools. Its target is subtle: an answer can contain a fact found in the evidence while crediting the wrong source. The system preserves tool-source identities, splits an answer into claims and checks both support and attribution after generation.
In the authors’ held-out evaluation, it caught 138 of 139 claims experts said should be blocked, while also holding 67 claims experts considered supported. A harder test with similar sources produced only 50.3% exact-source identification. These are reported research results from particular test sets, not proof that automated checking can replace an editor or domain expert.
Analysis
The trade-off is visible: catching more questionable claims can create extra review work. Keeping the specific evidence attached to each claim helps a reviewer understand why an answer passed or failed.
Action
When reviewing an AI answer, open the cited source rather than merely checking that a link exists. Ask whether that exact source supports the statement and its attribution—not just a related fact.
Apple Pay’s Indian launch begins with narrow coverage
Apple Pay has begun rolling out in India with Axis Bank, according to TechCrunch’s updated report. Initial support covers eligible Axis cards on Visa and Mastercard, not RuPay. Customers also need a merchant or terminal enabled for the service. The publication updated its earlier launch preview to say the rollout was underway.
The distinction from India’s UPI matters: Apple Pay is initially card-based, whereas UPI moves money between bank accounts. TechCrunch says HDFC Bank, ICICI Bank and SBI Card are not supporting the launch while commercial negotiations continue. It also reports that Apple and the banks did not respond to requests for comment, so details beyond observed availability retain that reporting limitation.
Analysis
This is an additional payment route, not an overnight replacement for an established national system. A supported card and an enabled terminal are separate requirements, which can make early acceptance uneven.
Action
If this rollout affects you, confirm your exact card’s eligibility and the merchant’s acceptance before relying on it. Keep an existing payment option available while coverage develops.
Humanoid walking research borrows from passive mechanics
A preprint submitted on 28 September explores training humanoid robots to walk more economically by borrowing ideas from passive dynamic walking. Early training uses a tilted-gravity field to encourage efficient movement, then removes that guidance before optimisation under ordinary dynamics. The authors say the framework requires no reference trajectories, gait phases or contact schedules.
On hardware, they report a 16.3% reduction in cost of transport when combined with a walking motion prior. Without that prior, the reported reduction was 4.5%, within trial-to-trial variation. That qualification is essential: the stronger result belongs to a combined setup. This is research reported by its authors, not a demonstrated battery-life improvement across commercial humanoids.
Analysis
The interesting idea is to guide discovery of a useful gait during training, rather than leave the robot dependent on that artificial assistance. The modest result without the motion prior also shows why experimental conditions belong beside the headline number.
Action
For robotics claims, check whether a result came from simulation or hardware, what other components were required, and how large the variation was. A percentage without those details is an incomplete comparison.