ANTHROPIC.COMClaude Opus 5
Introducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.
KNOWLEDGE REPOSITORY
AN IN-DEPTH REPOSITORY
Every model, paper, media release, agent and tool from the month—in one source-linked archive.
COMPLETE MONTHLY ARCHIVE
The frontier shipped: faster model families, natural voice, embodied navigation and the infrastructure needed to put agents into real work.
THE COMPLETE MONTHLY INDEX
Every linked paper, product release, research project and article has been checked directly. Cards use the publication's own description, imagery and published date wherever the source supplied them.
The AI Search helps discover the wider field, while every available card opens the original publication. Dates come only from publication metadata or official source APIs—never from a video. When no trustworthy source date exists, the card shows only its archive month.
ANTHROPIC.COMIntroducing Claude Opus 5
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.
BFL.AIFLUX 3: One Multi-Modal Model | Black Forest Labs
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
HUGGINGFACE.CObaseten/GLM-5.2-Vision-NVFP4 · Hugging Face
GLM 5.2 Vision NVFP4 is published on Hugging Face. Open the model or project card for the author’s documentation, files, license and usage details.
LEARN.CHATGPT.COMChatGPT Voice | ChatGPT Learn
Try voice in Chat, Work, and Codex in the ChatGPT desktop app.
Health in ChatGPT links to its original tools page, but the publisher did not expose a readable preview. Open the source directly for the full release.
HOMIE Demo Page
Human-object Centric Video Personalization via Multimodal Cognition Enhancement
POOLSIDE.AIIntroducing Laguna S 2.1
Today we’re releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.
HUGGINGFACE.COmicrosoft/Mage-Flow · Hugging Face
Mage-Flow is published on Hugging Face. Open the model or project card for the author’s documentation, files, license and usage details.
MODELSCOPE.AINanbeige4.2-3B
ModelScope——A one-stop service that brings together the most advanced machine learning models from various fields, offering model exploration experience, inference, training, deployment, and application.
OpenAI–Hugging Face security incident links to its original research page, but the publisher did not expose a readable preview. Open the source directly for the full release.
NEXT-STATE.GITHUB.IOHow to train a frontier-level world model
How we trained and open-sourced a frontier-level world model — the lessons, failures, and fixes — with a live, playable demo running on Reactor.
Qwen Studio
Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.
Hybrid Linear Attention with Attention Residuals for Efficient Video Generation.
RESEARCH.GOOGLETowards a quantum computer that learns from its errors
Volodymyr Sivak and Paul Klimov, Research Scientists, Google Quantum AI, Google Research
PENSIONER-11.GITHUB.IOShotPlan: Cinematic Video Generation with Learnable Planning Token
ShotPlan enables multi-shot cinematic video generation with frame-accurate hard cuts, soft transitions and temporally localized camera movement via learnable planning tokens.
HUGGINGFACE.COBringing Nunchaku 4-bit Diffusion Inference to Diffusers
Diffusers gained native loading for Nunchaku 4-bit diffusion checkpoints, reducing memory while accelerating denoising and removing the need for a separate inference engine.
RESEARCH.GOOGLESymptomAI: Towards a conversational AI agent for everyday symptom assessment
Google Research evaluated conversational symptom-assessment agents in a national-scale study with 13,917 participants, comparing their differential diagnoses with clinician assessments and wearable biosignals.
BLOG.GOOGLEIntroducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
HUGGINGFACE.COGrabette: an open system to record robot-manipulation data
Pollen Robotics open-sourced a handheld system for recording manipulation demonstrations and turning them into LeRobot datasets through a browser-based processing pipeline.
HUGGINGFACE.COIntroducing Cosmos 3 Edge
NVIDIA released a 4B open world model for on-device physical AI, combining visual understanding, prediction and robot-action generation for edge hardware including Jetson.
RESEARCH.NVIDIA.COMARDY: Autoregressive Diffusion for Interactive Motion
ARDY is an autoregressive diffusion model for interactive human motion generation with online text prompting and flexible kinematic constraints.
GENCEPTION.GITHUB.IOVideo Generation Models are General-Purpose Vision Learners
GenCeption: one unified, feed-forward vision model built on video generative pretraining — depth, normals, pose, segmentation, keypoints & 4D grounding with SOTA performance, steered by text instructions.
GITHUB.COMGitHub - google/GNM: An open ecosystem of parametric human models and perception stacks, starting with GNM Head.
An open ecosystem of parametric human models and perception stacks, starting with GNM Head. - google/GNM
KIMI.COMKimi K3 Tech Blog: Open Frontier Intelligence
Kimi K3 is the world's first open 3T-class model — frontier performance across coding, knowledge work, and reasoning, with native multimodality and 1M context.