Pathway’s Baby Dragon Hatchling (BDH) is a brain-inspired architecture that reasons in latent space. It differs from transformer architectures by not emitting chain-of-thought tokens. Development and scaling of BDH is occurring on Amazon SageMaker HyperPod. BDH-CQ achieved a new cost-efficiency mark on the ARC-AGI-1 benchmark. This indicates a potential improvement in efficiency compared to other models on that benchmark. This approach allows for potentially more efficient model execution, particularly for complex reasoning tasks. The use of SageMaker HyperPod provides the necessary compute resources for training and experimentation. This development represents a shift in architectural design for AI models, moving away from traditional token-based approaches.
Agents1 min read
Pathway BDH Development on SageMaker HyperPod
Pathway’s Baby Dragon Hatchling (BDH) architecture is being developed and scaled on Amazon SageMaker HyperPod. BDH-CQ achieved a new cost-efficiency mark on the ARC-AGI-1 benchmark.
By OpenSmartRoute editorial · written through the router by writer-small
From AWS machine learning blog - “Pathway’s brain-inspired architecture development on Amazon SageMaker HyperPod”

Keep reading
Related posts
Agents1 min read
SageMaker AI: G7, G6, and G5 LLM Inference Benchmarks
This benchmark compares the performance of Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B across G5, G6, G6e, and G7 GPU instances on SageMaker AI. G7 instances demonstrate price-performance gains for real-time LLM inference.
LLMs1 min read
Evidence integration in large language models analyzed through distributional theory
A distributional theory explains how large language models incorporate external evidence, revealing that model responses are influenced by prior beliefs and evidence characteristics across multiple domains.
Research1 min read
HarvestBench measures LLM agent decisions on animal harm in farm simulation
HarvestBench evaluates how language models decide to avoid harming animals in a simulated farm environment, with decisions priced and measured across multiple models and scenarios.
