UDIO.COMA New Era
A New Era of Music - Udio with Universal Music Group
Discover, create, and share music with the world. Use the latest technology to create AI music in seconds.

AN IN-DEPTH REPOSITORY
Every model, paper, media release, agent and tool from the month—in one source-linked archive.
COMPLETE MONTHLY ARCHIVE
Real-time video, multimodal image systems and practical agent platforms arrived alongside major releases from OpenAI, Google, Anthropic and the open-source field.
THE COMPLETE MONTHLY INDEX
Every linked paper, product release, research project and article has been checked directly. Cards use the publication's own description, imagery and published date wherever the source supplied them.
The AI Search helps discover the wider field, while every available card opens the original publication. Dates come only from publication metadata or official source APIs—never from a video. A checked date records our review, not the release. Older undated entries show their archive month.
UDIO.COMA New Era of Music - Udio with Universal Music Group
Discover, create, and share music with the world. Use the latest technology to create AI music in seconds.
GITHUB.COMGitHub - deepseek-ai/DeepSeek-OCR: Contexts Optical Compression
Contexts Optical Compression. Contribute to deepseek-ai/DeepSeek-OCR development by creating an account on GitHub.
G-VISTA.GITHUB.IOVISTA: A Test-Time Self-Improving Video Generation Agent
G Vista is linked to its original publication. Open the source for the author’s release notes, demonstrations and technical details.
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
Holo Cine is linked to its original publication. Open the source for the author’s release notes, demonstrations and technical details.
Inpaint4Drag
Inpaint4drag is linked to its original publication. Open the source for the author’s release notes, demonstrations and technical details.
Introducing Chatgpt Atlas is linked to its original publication. Open the source for the author’s release notes, demonstrations and technical details.
HUGGINGFACE.COkrea/krea-realtime-video · Hugging Face
Krea Realtime Video is published at its original model or project page. Open the source for documentation, files, demonstrations and usage details.
JAMESYJL.GITHUB.IONANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
Nano3D is linked to its original publication. Open the source for the author’s release notes, demonstrations and technical details.
GMARGO11.GITHUB.IOSoftMimic: Learning Compliant Whole-body Control from Examples
SoftMimic: Learning Compliant Whole-body Control from Examples
GITHUB.COMGitHub - vita-epfl/Stable-Video-Infinity: [ICLR 26 Oral] Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
[ICLR 26 Oral] Stable Video Infinity: Infinite-Length Video Generation with Error Recycling - vita-epfl/Stable-Video-Infinity
SJTUPLAYER.GITHUB.IOUltraGen: High-Resolution Video Generation with Hierarchical Attention
UltraGen: High-Resolution Video Generation with Hierarchical Attention
腾讯混元3D生成模型基于Diffusion技术,支持文本和图像生成3D资产。该模型配备精心设计的文本和图像编码器、扩散模型及3D解码器,能够实现多视图生成、重建及单视图生成。腾讯混元3D大模型可快速生成精美3D物体,适用于多种下游应用。
MISTRAL.AIIntroducing Mistral AI Studio
Mistral introduced an enterprise AI platform for building, evaluating and operating production applications with its models, observability and governance in one environment.
BLOG.GOOGLENew updates and more access to Google Earth AI
Earth AI is helping enterprises and cities with everything from environmental monitoring to disaster response.
BLOG.GOOGLEOur Quantum Echoes algorithm is a big step toward real-world applications for quantum computing
Our latest quantum breakthrough, Quantum Echoes, offers a path toward unprecedented scientific discoveries and analysis.
WORV-AI.GITHUB.IOD2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
Desktop gaming data effectively pretrains embodied AI: 152× compression via OWA Toolkit, YouTube pseudo-labeling with Generalist-IDM, achieving 96.6% on LIBERO manipulation and 83.3% on CANVAS navigation with 1.3K hours of data.
FENGHORA.GITHUB.IODiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
DiT360 Page is linked to its original publication. Open the source for the author’s release notes, demonstrations and technical details.
PBIHAO.GITHUB.IOIndex is linked to its original publication. Open the source for the author’s release notes, demonstrations and technical details.
FELIXTAUBNER.GITHUB.IOMVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
Our model can create realistic 4D avatars from a single reference image using multi-view video diffusion models.
NVIDIANEWS.NVIDIA.COMNVIDIA DGX Spark Arrives for World’s AI Developers
NVIDIA today announced it will start shipping NVIDIA DGX Spark™, the world’s smallest AI supercomputer.
PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System
A real-world generalizable and natural humanoid-scene interaction system.
KANGLIAO929.GITHUB.IOThinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
We make the first attempt to unify camera-centric understanding and generation in a cohesive multimodal framework.
HUGGINGFACE.COinclusionAI/Ring-1T · Hugging Face
Ring 1T is published at its original model or project page. Open the source for documentation, files, demonstrations and usage details.
WORLDLABS.AIRTFM: A Real-Time Frame Model
A research preview of RTFM, a new generative world model that generates video in real-time as you interact with it.