LLMs1 min read
ViCo: Visual-oriented Coding for Chart Replication with Self-Reflection
A new framework uses an 8B model to generate charts matching human visual standards through iterative self-reflection and reinforcement learning.
From arXiv cs.CL
Blog
Daily notes on new models, LLM releases, agent frameworks and AI research, written from the sources we follow and delivered as a newsletter every day.
Get the daily issue
Every new post of the day, in one email. Confirmation required.
LLMs1 min read
A new framework uses an 8B model to generate charts matching human visual standards through iterative self-reflection and reinforcement learning.
From arXiv cs.CL
AI1 min read
A new WhatsApp Business MCP server lets developers use AI coding agents like Claude, Cursor, Codex, and ChatGPT to handle setup, messaging templates, testing, and troubleshooting.
From TechCrunch AI
How this blog is made
Each feed entry becomes one request to OpenSmartRoute: the router picks a model with a cost-weighted objective, the editorial-writer skill is layered on the prompt, and the outcome trains the learners - the same pipeline available to every workspace.
Open any post to see which target answered, its confidence, the alternatives and what the request cost. Run the same pipeline yourself: register feeds in the operator console, map a small model under Providers, or call POST /api/v1/route with execute: true.
The research introduces Chopthin-Consensus Power Sampling (CCPS), a method for LLM decoding that preserves reasoning diversity and improves accuracy. Evaluation across open-weight models and benchmarks shows CCPS matches or exceeds the Power-SMC baseline in 14 of 15 settings, achieving gains of up to 10.6 percentage points.
From arXiv cs.CL
Research1 min read
Research identifies five axes of neural parameterization in recent Gaussian Splatting systems, revealing a nuanced approach beyond simple neural labels. A controlled study of mip-NeRF 360 demonstrates that sharing appearance and opacity improves reconstruction quality, while geometric structure decoding provides no further benefit.
From arXiv cs.AI
Research1 min read
A study comparing Claude-agent-sdk and deepagents across Claude-opus-4-8 and openai-codex with gpt-5.5, gemini-3.5-flash and deepseek-v3.2 found no significant advantage for either vendor-native harness. The study also examined harness cost and completion rates, revealing higher costs and a complex, unresolved billing structure.
From arXiv cs.AI
LLMs1 min read
This web app, commit-rewriter, edits commit messages for repositories, specifically designed to remove coding agent cruft and issue ID references from Datasette security releases. It creates a timestamped branch for reverting edits and rewrites commits from the first edited to the most recent.
From Simon Willison
Agents1 min read
Credit Genie utilizes OpenWiki to automatically maintain and update codebase documentation, reducing reliance on individual knowledge and providing searchable context for engineering models and agents.
From LangChain blog
AI1 min read
Cognition’s SWE-2, a post-trained model based on Kimi K3, achieves 50% accuracy on FrontierCode 1.1 Main, costing 64% less than Fable 5.1. It’s currently available only within the Devin platform.
From MarkTechPost
LLMs1 min read
A new LLM pipeline using Hyper-Parallel Decoding achieves high extraction accuracy and significant cost reduction for e-commerce attribute extraction. This enables production-scale use for building structured product knowledge bases.
From arXiv cs.CL
LLMs1 min read
Osprey pretrains a small language model to accelerate speculative decoding, improving acceptance rates across diverse target models. The system reduces per-target adaptation work through a reusable backbone and lightweight adjustments.
From arXiv cs.CL
LLMs1 min read
Simon Willison created a tool to view Blender models directly in the browser using a .blend URL. The tool was built by prompting GPT-6 Astra with a ChatGPT Images 2.5 generated image and using Codex to create a Blender model from it.
From Simon Willison
Agents1 min read
This post details deploying the 2.4-trillion parameter Qwen3.8-2.4T-A95B model on Amazon SageMaker HyperPod using vLLM. The setup includes NVFP4 quantization and an OpenAI-compatible endpoint with reasoning and tool calling capabilities.
From AWS machine learning blog
Research1 min read
PGP-Clinical-TimeKAN forecasts clinical trajectories using a novel framework combining missingness-aware encoding and a soft organ-system prior. The model achieves state-of-the-art normalized MAE and RMSE on MIMIC-IV data, demonstrating the value of joint trajectory forecasting.
From arXiv cs.AI
Research1 min read
A new design approach, TEAM-Design, optimizes human-agent team deployments by strategically allocating replay budgets to the most uncertain comparisons. This reduces wasted time and compute when human-AI workflows don't outperform individual alternatives.
From arXiv cs.AI
LLMs1 min read
A new inference method, Intra-Prompt Parallel Decoding (IPPD), achieves up to 7x throughput in common-context question answering by decoding multiple questions within a single prompt. This approach overcomes GPU memory bottlenecks and outperforms existing techniques like prefix caching.
From arXiv cs.CL
LLMs1 min read
Tool: Mercator ↔ Equal Earth I got curious about the Equal Earth map projection that was recently voted on at the UN so I had GPT-6 Astra (medium) in ChatGPT Work build me this animated transition between Mercator and Equal Earth using D...
From Simon Willison
LLMs1 min read
A 0.6B language model's behavior was examined, revealing calibration failures where internal verdicts are misaligned with output logits, affecting accuracy and interpretability.
From arXiv cs.CL
LLMs1 min read
A new evaluation scheme based on the Hüllermeier-Rifqi Index assesses phonetic encoding algorithms' conformity to word-based transcriptions in IPA, using normalized edit distance and a random string adjustment.
From arXiv cs.CL
Research1 min read
The $ au^ au$-bench evaluates agent building from real business data, requirements, and APIs, measuring performance across multiple tasks to reflect real client engagement conditions.
From arXiv cs.AI
Posts are drafted from public feeds by models OpenSmartRoute routes to - the same router, skill and metering customers use - and always link to the original source. Corrections: support.
Archive (34)