Modernizing Table Batched Embeddings with FBTriton
PyTorchby Meta Team: Daohang Shi, Oleksandr Stashuk, Rupert Wu, Liangbei Xu, Rich Zhu·1d ago
Featured projects This post explores the FBTriton kernel design for Table Batched Embedding (TBE) forward and backward passes.
Evolution of the PyTorch Media Processing Landscape
PyTorchby Nicolas Hug (Meta), Scott Schneider (Meta)·2d ago
Featured projects TL;DR If you need to decode or encode media, whether it’s images, video, or audio, use TorchCodec.
PyTorch Hardware Enablement: Updates from the Accelerator Integration Working Group
PyTorchby PyTorch TAC Accelerator Integration Working Group: Anisha Kushwaha, Atharva Kshirsagar, Jiahao Chen, Jiahao Tan, Jiawei Li, Jewel K. M., Mansi Agarwal, Parshant Sharma, Riya Punia, Subin George, Tanma·3d ago
Featured projects TL;DR The Accelerator Integration Working Group plays a vital role in standardizing how new hardware architectures connect to the open source AI ecosystem.
Building a High-Performance and Portable vLLM Linear Backend with Helion
PyTorchby Sean Chen (Red Hat) and Shangdi Yu (PyTorch, Meta Platforms)·6d ago
Featured projects TL;DR We integrated Helion into vLLM’s linear backend to explore how an autotuned, high-level kernel DSL can improve LLM inference performance while reducing kernel implementation...
New Pathway to PyTorch Certified Associate (PTCA) Certification
PyTorchby PyTorch Foundation·6d ago
We’re excited to bring you the new PyTorch Certified Associate (PTCA) Certification Pathway, combining focused learning modules and the PyTorch Certified Associate (PTCA) certification exam...
Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell
PyTorchby Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra), Darren Liu, Dev (Devashish) Shankar, Max Leung, Nick Riasanovsky, Hao Yan, Manman Ren, Yuanwei (Kevin) Fang·6d ago
Featured projects TL;DR In this blog post, we present our work on Jagged Flash Attention (JFA) — the attention kernel behind Meta’s Generative Ads Model (GEM) — on NVIDIA Blackwell (B200), built...
A Ray-Focused Guide to PyTorch Conference North America
PyTorchby PyTorch Foundation·8d ago
Featured projects TL;DR In only a few weeks, PyTorch Conference North America 2026 will begin in San Jose, bringing together the open source AI community to share ideas and collaborate.
From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR
PyTorchby Mehant Kammakomati (IBM), Jewel K M (Red Hat), Anubhav Jana (IBM), Padmanabha Venkatagiri Seshadri (IBM)·8d ago
Featured projects TL;DR PyTorch’s Cross-Repository CI Relay (CRCR) gives out-of-tree accelerators a clean, scalable way to plug into upstream CI – and it leaves each backend free to decide which of...
Accelerate Your AI Journey with new Introduction Track at PyTorch Conference NA 2026 and PyTorch Associate Training
PyTorchby PyTorch Foundation·13d ago
Featured projects As deep learning models move rapidly from research prototypes into core enterprise infrastructure, the demand for practical, end-to-end PyTorch expertise has never been higher.
From Research Project to Open Source Ecosystem: Bring Your Academic PyTorch Project to PyTorchCon NA
PyTorchby PyTorch Foundation·15d ago
Some of the most interesting work being built with PyTorch starts in universities, research labs, student groups, and academic institutions. A new model architecture.
Hardware-Agnostic Models in vLLM
PyTorchby Thomas Parnell (IBM), Thomas Ortner (IBM), Richard Zou (Meta), Harry Mellor (Hugging Face)·16d ago
Featured projects TL;DR To achieve state-of-the-art performance at the frontier, vLLM is changing its internal implementation in ways that make it incompatible with fullgraph torch.compile.
How Shopify built a continual learning loop with PyTorch and vLLM
PyTorchby Cody Mazza-Anthony, Sr. Staff Machine Learning Engineer, @cmazzaanthony & Andrew McNamara, VP Machine Learning, @drewch·16d ago
Featured projects TL;DR: This case study explores how Shopify compresses production failures into model weights every day, beats frontier-model quality, and cuts serving costs 96% by building a...
TinyTorch: Don’t Just Import PyTorch. Build It.
PyTorchby Vijay Janapa Reddi, Harvard University and ETH Zurich · Andrea Mattia Garavagno, ETH Zurich·16d ago
Featured projects A framework you write yourself, tensors through transformers TL;DR Every mature systems project eventually needs a teaching version.
PyTorch Day Japan 2026 Comes to Tokyo on December 10
PyTorchby PyTorch Foundation·20d ago
PyTorch Day Japan 2026 will bring the open source AI community together in Tokyo on December 10 for a full day of technical talks and interactive discussions designed to foster knowledge exchange...
Open Research, Tooling & Optimization at PyTorch Conference North America 2026
PyTorchby PyTorch Foundation·22d ago
Featured projects TL;DR Taking place October 20 to 21 in San Jose, California, PyTorch Conference North America 2026 highlights open research, tooling, and performance optimization across compiler...
Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell
PyTorchby Dev (Devashish) Shankar, Darren Liu, Chunzhi Yang, Jackie (Jiaqi) Xu,Markus Hoehnerbach, Jason Xie, Santosh Mohan, Han Xu, Rich Zhu, Josh Fromm, Hongtao Yu, Max Leung, and John Bocharov·22d ago
TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes.
Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels
PyTorchby Sayak Paul (Hugging Face), Dunfan Lu (Meta), Tarindu Jayatilaka (Meta), Jongsok Choi (Meta)·27d ago
Featured projects TL;DR The HuggingFace Kernels project now has Helion support. This blog walks through how to build, autotune, and ship performant and portable Helion kernels via the Hugging Face...
PyTorch Conference China 2026: Advancing the Open Source AI Stack
PyTorchby PyTorch Foundation·28d ago
PyTorch Conference China 2026 brought the PyTorch community together in Shanghai on September 8–9 alongside KubeCon + CloudNativeCon and OpenInfra Summit, following sponsor-hosted co-located events...
Alibaba Cloud, Ant Group, Cambricon and Huawei Come Together in Shanghai to Advance the Open Source AI Stack at PyTorch Conference China
PyTorchby PyTorch Foundation·8 Sept 2026
China’s leading AI technology companies Alibaba Cloud, Cambricon and Ant Group join the PyTorch Foundation as members Summary - Alibaba Cloud and Cambricon joined the PyTorch Foundation as Platinum...
Cambricon Joins the PyTorch Foundation as a Platinum Member
PyTorchby PyTorch Foundation·8 Sept 2026
The PyTorch Foundation, a community-driven hub for open source AI under the Linux Foundation, is announcing today that Cambricon has joined as a Platinum member.
Your filters hide everything on this page. Adjust them in preferences.