LLMs1 min read
BenchMIRT analyzes what LLM benchmarks actually measure
BenchMIRT investigates the alignment between LLM benchmark tasks and real-world capabilities, highlighting potential discrepancies for model evaluation.
From Hugging Face blog
Blog
Daily notes on new models, LLM releases, agent frameworks and AI research, written from the sources we follow and delivered as a newsletter every day.
Get the daily issue
Every new post of the day, in one email. Confirmation required.
LLMs1 min read
BenchMIRT investigates the alignment between LLM benchmark tasks and real-world capabilities, highlighting potential discrepancies for model evaluation.
From Hugging Face blog
Agents1 min read
Z.ai released GLM-5.3-Flash, a natively multimodal model with a 1M-token context window and 320B parameters, achieving strong performance benchmarks and competitive pricing, sparking significant community interest.
From Latent Space
How this blog is made
Each feed entry becomes one request to OpenSmartRoute: the router picks a model with a cost-weighted objective, the editorial-writer skill is layered on the prompt, and the outcome trains the learners - the same pipeline available to every workspace.
Open any post to see which target answered, its confidence, the alternatives and what the request cost. Run the same pipeline yourself: register feeds in the operator console, map a small model under Providers, or call POST /api/v1/route with execute: true.
Gemini 3.5 Transcribe offers enhanced speech-to-text transcription capabilities, aiming to improve accuracy and intelligence in transcription tasks.
From Google DeepMind blog
LLMs1 min read
Alibaba has released the open weights for Qwen3.8-2.4T-A95B, a 2.4 trillion parameter model, allowing near-frontier capabilities to be deployed on NVIDIA GB300 NVL72 systems. This enables engineers to run large language models with configurable reasoning.
From NVIDIA technical blog
Research2 min read
Google Research introduces Mobility-Embedded POIs (ME-POIs), a framework that integrates mobility patterns with language models to improve predictions about place attributes like operating hours and busyness. This approach addresses data sparsity and enhances model accuracy.
From Google Research blog
LLMs1 min read
This tutorial demonstrates building a basic AI text detector from scratch, utilizing a DistilBERT classifier and RLVR. The project aims to illustrate AI detector functionality and serve as a case study for building scoring systems alongside a user-friendly UI.
From Ahead of AI (Sebastian Raschka)
LLMs1 min read
This report details the evolving state of open models, focusing on size trends, licensing options, and key performance indicators for models deployed in production environments. It highlights shifts in model architecture and accessibility for engineers.
From Hugging Face blog
Research1 min read
Gemini 3.7 Flash is a new language model introduced by Google DeepMind, designed for improved performance and efficiency in AI applications.
From Google DeepMind blog
Research1 min read
Microsoft Research introduced MindTopo, a benchmark designed to assess a VLM's ability to understand topological relationships like paths and knots. This tool provides a new method for evaluating and improving spatial reasoning and planning capabilities in AI models.
From Microsoft Research
Research2 min read
Research indicates that recall failures, not encoding limitations, are the primary cause of factual errors in advanced LLMs like Gemini-3 and GPT-5. The new knowledge profiling framework highlights this issue and suggests inference-time methods as a key area for improvement.
From Google Research blog
LLMs1 min read
NVIDIA’s Nemotron 3.5 Lightning, a 30 billion parameter open model, is now accessible on Ollama for local agent execution. Optimized for multi-step tasks and long context windows, it offers 4x higher throughput and faster task completion compared to similar models.
From Ollama blog
Research1 min read
Google DeepMind and A24 have initiated a research partnership focused on developing advanced AI agents. The collaboration aims to explore the use of large language models in creative workflows, specifically for scriptwriting.
From Google DeepMind blog
Research1 min read
Google DeepMind has released Nano Banana 2 Lite, a smaller language model, and Gemini Omni Flash, a new flash storage solution optimized for AI workloads. These tools are designed to enable efficient and cost-effective AI development and deployment.
From Google DeepMind blog
Research1 min read
Google DeepMind has introduced the ability for Gemini 3.5 Flash to utilize computer tools, expanding its capabilities beyond traditional language model tasks. This allows the model to interact with external applications and services, enhancing its utility for a wider range of use cases.
From Google DeepMind blog
Posts are drafted from public feeds by models OpenSmartRoute routes to - the same router, skill and metering customers use - and always link to the original source. Corrections: support.
Archive (34)