Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion async-rl-training-landscape.md
Original file line number Diff line number Diff line change
Expand Up @@ -207,7 +207,7 @@ The table above reveals a striking pattern: **eight of the sixteen libraries sur

The cost is a hard dependency on a non-trivial runtime. This trade-off can be worthwhile, especially for production-scale training (64+ GPUs, multi-day runs, complicated reward computation).

While Ray's actor model is the main player on the field, [Monarch](https://github.com/pytorch/monarch) emerged as a new PyTorch-native distributed actor framework from Meta, purpose-built for GPU workloads. Like Ray, Monarch is based on the actor model; components are independent actors with mailboxes communicating via messages, but it is designed from the ground up for the PyTorch/CUDA ecosystem rather than being a general-purpose distributed runtime.
While Ray's actor model is the main player on the field, [Monarch](https://github.com/meta-pytorch/monarch) emerged as a new PyTorch-native distributed actor framework from Meta, purpose-built for GPU workloads. Like Ray, Monarch is based on the actor model; components are independent actors with mailboxes communicating via messages, but it is designed from the ground up for the PyTorch/CUDA ecosystem rather than being a general-purpose distributed runtime.

Monarch offers several capabilities particularly relevant to async RL. An [example implementation of async RL with Monarch](https://allenwang28.github.io/monarch-gpu-mode/05_rl_intro.html) (from the GPU Mode lecture series) demonstrates the architecture: generators, a replay buffer, and a trainer are modelled as Monarch actors, with the replay buffer absorbing latency variance from straggler rollouts and RDMA weight sync pushing updated parameters to generators without blocking training. The pattern is structurally identical to Ray-based designs (verl, SkyRL, open-instruct) but implemented with pure PyTorch-native primitives.

Expand Down
6 changes: 3 additions & 3 deletions build-rocm-kernels.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ authors:

## Intoduction

Custom kernels are the backbone of high-performance deep learning, enabling GPU operations tailored precisely to your workload; whether that’s image processing, tensor transformations, or other compute-heavy tasks. But compiling these kernels for the right architectures, wiring all the build flags, and integrating them cleanly into PyTorch extensions can quickly become a mess of CMake/Nix, compiler errors, and ABI issues, which is not fun. Hugging Face’s [**kernels**](https://github.com/huggingface/kernels) library makes it easy to build (with [**kernel-builder**](https://github.com/huggingface/kernels/tree/main/builder)) and share these kernels with the [**kernels-community**](https://huggingface.co/kernels-community), with support for multiple GPU and accelerator backends, including CUDA, ROCm, Metal, and XPU. This ensures your kernels are fast, portable, and seamlessly integrated with PyTorch.
Custom kernels are the backbone of high-performance deep learning, enabling GPU operations tailored precisely to your workload; whether that’s image processing, tensor transformations, or other compute-heavy tasks. But compiling these kernels for the right architectures, wiring all the build flags, and integrating them cleanly into PyTorch extensions can quickly become a mess of CMake/Nix, compiler errors, and ABI issues, which is not fun. Hugging Face’s [**kernels**](https://github.com/huggingface/kernels) library makes it easy to build (with [**kernel-builder**](https://github.com/huggingface/kernels/tree/main/kernel-builder)) and share these kernels with the [**kernels-community**](https://huggingface.co/kernels-community), with support for multiple GPU and accelerator backends, including CUDA, ROCm, Metal, and XPU. This ensures your kernels are fast, portable, and seamlessly integrated with PyTorch.

In this guide, we focus exclusively on ROCm-compatible kernels and show how to build, test, and share them using [kernels](https://github.com/huggingface/kernels/tree/main). You’ll learn how to create kernels that run efficiently on AMD GPUs, along with best practices for reproducibility, packaging, and deployment.

Expand Down Expand Up @@ -110,7 +110,7 @@ If you look at the original files of the gemm kernel in the RadeonFlow Kernels,
- Use `.h` for header files containing kernel declarations, inline functions, or template code that will be included in other files
- Use `.hip` for implementation files containing HIP/GPU code that needs to be compiled separately (e.g., kernel launchers, device functions with complex implementations)

In our example, `gemm_kernel.h`, `gemm_kernel_legacy.h`, and `transpose_kernel.h` are header files, while `gemm_launcher.hip` is a HIP implementation file. This naming convention helps the kernel-builder ([`kernels/builder`](https://github.com/huggingface/kernels/tree/main/builder)) correctly identify and compile each file type.
In our example, `gemm_kernel.h`, `gemm_kernel_legacy.h`, and `transpose_kernel.h` are header files, while `gemm_launcher.hip` is a HIP implementation file. This naming convention helps the kernel-builder ([`kernels/builder`](https://github.com/huggingface/kernels/tree/main/kernel-builder)) correctly identify and compile each file type.

### Step 2: Configuration Files Setup

Expand Down Expand Up @@ -526,7 +526,7 @@ Building and sharing ROCm kernels with the Hugging Face is now easier than ever.

## Related Libraries & Hub

- [kernels](https://github.com/huggingface/kernels) – Library to build, manage and load kernels from the Hub. It contains the [kernel-builder](https://github.com/huggingface/kernels/tree/main/builder) tooling.
- [kernels](https://github.com/huggingface/kernels) – Library to build, manage and load kernels from the Hub. It contains the [kernel-builder](https://github.com/huggingface/kernels/tree/main/kernel-builder) tooling.
- [Kernels Community Hub](https://huggingface.co/kernels-community) – Share and discover kernels from the community.


6 changes: 3 additions & 3 deletions custom-cuda-kernels-agent-skills.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,7 @@ Build an optimized attention kernel for H100 targeting the Qwen3-8B model in tra

The agent can read the skill, select the right architecture parameters, generate the CUDA source, write the PyTorch bindings, set up `build.toml`, and create a benchmark script.

If you're working on more complex kernels, or architecture-specific optimizations, that aren't covered in the skill, then the skill supplies the fundamental building blocks and patterns to get you started. We are also open to contributions on the [skill itself](https://github.com/huggingface/kernels/tree/main/.docs/skills).
If you're working on more complex kernels, or architecture-specific optimizations, that aren't covered in the skill, then the skill supplies the fundamental building blocks and patterns to get you started. We are also open to contributions on the [skill itself](https://github.com/huggingface/kernels/tree/main/kernel-builder/skills).

## What is in the skill

Expand Down Expand Up @@ -249,7 +249,7 @@ cuda-capabilities = ["9.0"] # H100

### 2. Build all variants with Nix

Kernel Hub kernels must support all recent PyTorch and CUDA configurations. The kernel-builder Nix flake handles this automatically. Copy the [example `flake.nix`](https://github.com/huggingface/kernels/blob/main/builder/examples/relu/flake.nix) into your project and run:
Kernel Hub kernels must support all recent PyTorch and CUDA configurations. The kernel-builder Nix flake handles this automatically. Copy the [example `flake.nix`](https://github.com/huggingface/kernels/blob/main/examples/kernels/relu/flake.nix) into your project and run:

```shell
nix flake update
Expand Down Expand Up @@ -291,7 +291,7 @@ We built an agent skill that teaches coding agents how to write production CUDA

## Resources

- [CUDA Kernels Skill in `kernels`](https://github.com/huggingface/kernels/tree/main/skills/cuda-kernels)
- [CUDA Kernels Skill in `kernels`](https://github.com/huggingface/kernels/tree/main/kernel-builder/skills/cuda-kernels)
- [HuggingFace Kernel Hub Blog](https://huggingface.co/blog/hello-hf-kernels)
- [We Got Claude to Fine-Tune an Open Source LLM](https://huggingface.co/blog/hf-skills-training)
- [We Got Claude to Teach Open Models](https://huggingface.co/blog/upskill)
Expand Down
2 changes: 1 addition & 1 deletion eee-community-evals.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,7 @@ Hugging Face stores eval scores in the model repo as a YAML under `.eval_results

Submit your full records to [the EEE datastore](https://huggingface.co/datasets/evaleval/EEE_datastore).

Utilizing EEE requires only one additional step, which the converter largely automates. The [community eval converter tool](https://github.com/evaleval/every_eval_ever/tree/main/tools/hf-community-evals) can be found in the GitHub repository. To process a collection, execute the following:
Utilizing EEE requires only one additional step, which the converter largely automates. The [community eval converter tool](https://github.com/evaleval/every_eval_ever/blob/main/every_eval_ever/tools/hf_community_evals.py) can be found in the GitHub repository. To process a collection, execute the following:

```shell
uv run tools/hf-community-evals/community_evals_converter.py MMLU-Pro \
Expand Down
2 changes: 1 addition & 1 deletion gemma4.md
Original file line number Diff line number Diff line change
Expand Up @@ -633,7 +633,7 @@ mlx_vlm.generate \
--kv-quant-scheme turboquant
```

For audio examples and more details, please check [the MLX collection](https://hf.co/mlx-community/gemma-4).
For audio examples and more details, please check [the MLX collection](https://huggingface.co/collections/mlx-community/gemma-4).

### Mistral.rs

Expand Down
2 changes: 1 addition & 1 deletion moe-transformers.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,7 +130,7 @@ to:

### Dynamic Weight Loading with `WeightConverter`

The central abstraction introduced by this refactor is **dynamic weight loading** via a [`WeightConverter`](https://huggingface.co/docs/transformers/main/en/internal/weight_converter).
The central abstraction introduced by this refactor is **dynamic weight loading** via a [`WeightConverter`](https://huggingface.co/docs/transformers/main/en/weightconverter).

`WeightConverter` lets us define:

Expand Down
2 changes: 1 addition & 1 deletion nvidia-reachy-mini.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ We’ll be using the following:
1. A reasoning model: demo uses [NVIDIA Nemotron 3 Nano](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16)
2. A vision model: demo uses [NVIDIA Nemotron Nano 2 VL](https://huggingface.co/nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16)
3. A text-to-speech model: demo uses [ElevenLabs](https://elevenlabs.io)
4. [Reachy Mini](https://www.pollen-robotics.com/reachy-mini/) (or [Reachy Mini Simulation](https://github.com/pollen-robotics/reachy_mini/blob/develop/docs/platforms/simulation/get_started.md))
4. [Reachy Mini](https://www.pollen-robotics.com/reachy-mini/) (or [Reachy Mini Simulation](https://github.com/pollen-robotics/reachy_mini/blob/main/docs/source/platforms/simulation/get_started.md))
5. Python v3.10+ environment, with [uv](https://docs.astral.sh/uv/)

Feel free to adapt the recipe and make it your own \- you have many ways to integrate the models into your application:
Expand Down
2 changes: 1 addition & 1 deletion revamped-kernels.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,4 +170,4 @@ To solve this issue, kernels now link `libstdc++` dynamically. To ensure compati

Our goal with the Kernels project is to serve both kernel developers and users of custom kernels. We’re always keen on receiving feedback from the community on how we can improve it. Don’t hesitate to contribute!

*Acknowledgements: Thanks to [Aritra](ariG23498) for reviewing the post.*
*Acknowledgements: Thanks to [Aritra](https://huggingface.co/ariG23498) for reviewing the post.*
8 changes: 4 additions & 4 deletions unsloth-jobs.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,14 +96,14 @@ Codex discovers skills through [`AGENTS.md`](https://developers.openai.com/codex
Install individual skills with `$skill-installer`:

```text
$skill-installer install https://github.com/huggingface/skills/tree/main/skills/hugging-face-model-trainer
$skill-installer install https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer
```

For more details, see the [Codex Skills docs](https://developers.openai.com/codex/skills) and the [AGENTS.md guide](https://developers.openai.com/codex/guides/agents-md).

### Anything else

A generic install method is simply to clone the [skills repository](https://github.com/huggingface/skills) and copy the [skill](https://github.com/huggingface/skills/tree/main/skills/hugging-face-model-trainer) to your agent's skills directory.
A generic install method is simply to clone the [skills repository](https://github.com/huggingface/skills) and copy the [skill](https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer) to your agent's skills directory.

```text
git clone https://github.com/huggingface/skills.git
Expand All @@ -118,7 +118,7 @@ Once the skill is installed, ask your coding agent to train a model:
Train LiquidAI/LFM2.5-1.2B-Instruct on mlabonne/FineTome-100k using Unsloth on HF Jobs
```

The agent will generate a training script based on an [example in the skill](https://github.com/huggingface/skills/blob/main/skills/hugging-face-model-trainer/scripts/unsloth_sft_example.py), submit the training to HF Jobs, and provide a monitoring link via Trackio.
The agent will generate a training script based on an [example in the skill](https://github.com/huggingface/skills/blob/main/skills/huggingface-llm-trainer/scripts/unsloth_sft_example.py), submit the training to HF Jobs, and provide a monitoring link via Trackio.

## How It Works

Expand All @@ -131,7 +131,7 @@ Training jobs run on [Hugging Face Jobs](https://huggingface.co/docs/huggingface

### Example Training Script

The skill generates scripts like this based on the example in the [skill](https://github.com/huggingface/skills/blob/main/skills/hugging-face-model-trainer/scripts/unsloth_sft_example.py).
The skill generates scripts like this based on the example in the [skill](https://github.com/huggingface/skills/blob/main/skills/huggingface-llm-trainer/scripts/unsloth_sft_example.py).

```python
# /// script
Expand Down
2 changes: 1 addition & 1 deletion upskill.md
Original file line number Diff line number Diff line change
Expand Up @@ -322,4 +322,4 @@ The approach works for any specialized task where you'd otherwise write detailed

- [Upskill repo](https://github.com/huggingface/upskill)
- [Agent Skills Specification](https://agentskills.io)
- [HuggingFace kernel-builder](https://github.com/huggingface/kernels/tree/main/builder)
- [HuggingFace kernel-builder](https://github.com/huggingface/kernels/tree/main/kernel-builder)