diff --git a/async-rl-training-landscape.md b/async-rl-training-landscape.md index 7fa2858c4c..3c87479f96 100644 --- a/async-rl-training-landscape.md +++ b/async-rl-training-landscape.md @@ -207,7 +207,7 @@ The table above reveals a striking pattern: **eight of the sixteen libraries sur The cost is a hard dependency on a non-trivial runtime. This trade-off can be worthwhile, especially for production-scale training (64+ GPUs, multi-day runs, complicated reward computation). -While Ray's actor model is the main player on the field, [Monarch](https://github.com/pytorch/monarch) emerged as a new PyTorch-native distributed actor framework from Meta, purpose-built for GPU workloads. Like Ray, Monarch is based on the actor model; components are independent actors with mailboxes communicating via messages, but it is designed from the ground up for the PyTorch/CUDA ecosystem rather than being a general-purpose distributed runtime. +While Ray's actor model is the main player on the field, [Monarch](https://github.com/meta-pytorch/monarch) emerged as a new PyTorch-native distributed actor framework from Meta, purpose-built for GPU workloads. Like Ray, Monarch is based on the actor model; components are independent actors with mailboxes communicating via messages, but it is designed from the ground up for the PyTorch/CUDA ecosystem rather than being a general-purpose distributed runtime. Monarch offers several capabilities particularly relevant to async RL. An [example implementation of async RL with Monarch](https://allenwang28.github.io/monarch-gpu-mode/05_rl_intro.html) (from the GPU Mode lecture series) demonstrates the architecture: generators, a replay buffer, and a trainer are modelled as Monarch actors, with the replay buffer absorbing latency variance from straggler rollouts and RDMA weight sync pushing updated parameters to generators without blocking training. The pattern is structurally identical to Ray-based designs (verl, SkyRL, open-instruct) but implemented with pure PyTorch-native primitives. diff --git a/build-rocm-kernels.md b/build-rocm-kernels.md index b7a175bce0..558005d7b5 100644 --- a/build-rocm-kernels.md +++ b/build-rocm-kernels.md @@ -12,7 +12,7 @@ authors: ## Intoduction -Custom kernels are the backbone of high-performance deep learning, enabling GPU operations tailored precisely to your workload; whether that’s image processing, tensor transformations, or other compute-heavy tasks. But compiling these kernels for the right architectures, wiring all the build flags, and integrating them cleanly into PyTorch extensions can quickly become a mess of CMake/Nix, compiler errors, and ABI issues, which is not fun. Hugging Face’s [**kernels**](https://github.com/huggingface/kernels) library makes it easy to build (with [**kernel-builder**](https://github.com/huggingface/kernels/tree/main/builder)) and share these kernels with the [**kernels-community**](https://huggingface.co/kernels-community), with support for multiple GPU and accelerator backends, including CUDA, ROCm, Metal, and XPU. This ensures your kernels are fast, portable, and seamlessly integrated with PyTorch. +Custom kernels are the backbone of high-performance deep learning, enabling GPU operations tailored precisely to your workload; whether that’s image processing, tensor transformations, or other compute-heavy tasks. But compiling these kernels for the right architectures, wiring all the build flags, and integrating them cleanly into PyTorch extensions can quickly become a mess of CMake/Nix, compiler errors, and ABI issues, which is not fun. Hugging Face’s [**kernels**](https://github.com/huggingface/kernels) library makes it easy to build (with [**kernel-builder**](https://github.com/huggingface/kernels/tree/main/kernel-builder)) and share these kernels with the [**kernels-community**](https://huggingface.co/kernels-community), with support for multiple GPU and accelerator backends, including CUDA, ROCm, Metal, and XPU. This ensures your kernels are fast, portable, and seamlessly integrated with PyTorch. In this guide, we focus exclusively on ROCm-compatible kernels and show how to build, test, and share them using [kernels](https://github.com/huggingface/kernels/tree/main). You’ll learn how to create kernels that run efficiently on AMD GPUs, along with best practices for reproducibility, packaging, and deployment. @@ -110,7 +110,7 @@ If you look at the original files of the gemm kernel in the RadeonFlow Kernels, - Use `.h` for header files containing kernel declarations, inline functions, or template code that will be included in other files - Use `.hip` for implementation files containing HIP/GPU code that needs to be compiled separately (e.g., kernel launchers, device functions with complex implementations) -In our example, `gemm_kernel.h`, `gemm_kernel_legacy.h`, and `transpose_kernel.h` are header files, while `gemm_launcher.hip` is a HIP implementation file. This naming convention helps the kernel-builder ([`kernels/builder`](https://github.com/huggingface/kernels/tree/main/builder)) correctly identify and compile each file type. +In our example, `gemm_kernel.h`, `gemm_kernel_legacy.h`, and `transpose_kernel.h` are header files, while `gemm_launcher.hip` is a HIP implementation file. This naming convention helps the kernel-builder ([`kernels/builder`](https://github.com/huggingface/kernels/tree/main/kernel-builder)) correctly identify and compile each file type. ### Step 2: Configuration Files Setup @@ -526,7 +526,7 @@ Building and sharing ROCm kernels with the Hugging Face is now easier than ever. ## Related Libraries & Hub -- [kernels](https://github.com/huggingface/kernels) – Library to build, manage and load kernels from the Hub. It contains the [kernel-builder](https://github.com/huggingface/kernels/tree/main/builder) tooling. +- [kernels](https://github.com/huggingface/kernels) – Library to build, manage and load kernels from the Hub. It contains the [kernel-builder](https://github.com/huggingface/kernels/tree/main/kernel-builder) tooling. - [Kernels Community Hub](https://huggingface.co/kernels-community) – Share and discover kernels from the community. diff --git a/custom-cuda-kernels-agent-skills.md b/custom-cuda-kernels-agent-skills.md index e9405b6f11..8f7d15fbf3 100644 --- a/custom-cuda-kernels-agent-skills.md +++ b/custom-cuda-kernels-agent-skills.md @@ -77,7 +77,7 @@ Build an optimized attention kernel for H100 targeting the Qwen3-8B model in tra The agent can read the skill, select the right architecture parameters, generate the CUDA source, write the PyTorch bindings, set up `build.toml`, and create a benchmark script. -If you're working on more complex kernels, or architecture-specific optimizations, that aren't covered in the skill, then the skill supplies the fundamental building blocks and patterns to get you started. We are also open to contributions on the [skill itself](https://github.com/huggingface/kernels/tree/main/.docs/skills). +If you're working on more complex kernels, or architecture-specific optimizations, that aren't covered in the skill, then the skill supplies the fundamental building blocks and patterns to get you started. We are also open to contributions on the [skill itself](https://github.com/huggingface/kernels/tree/main/kernel-builder/skills). ## What is in the skill @@ -249,7 +249,7 @@ cuda-capabilities = ["9.0"] # H100 ### 2. Build all variants with Nix -Kernel Hub kernels must support all recent PyTorch and CUDA configurations. The kernel-builder Nix flake handles this automatically. Copy the [example `flake.nix`](https://github.com/huggingface/kernels/blob/main/builder/examples/relu/flake.nix) into your project and run: +Kernel Hub kernels must support all recent PyTorch and CUDA configurations. The kernel-builder Nix flake handles this automatically. Copy the [example `flake.nix`](https://github.com/huggingface/kernels/blob/main/examples/kernels/relu/flake.nix) into your project and run: ```shell nix flake update @@ -291,7 +291,7 @@ We built an agent skill that teaches coding agents how to write production CUDA ## Resources -- [CUDA Kernels Skill in `kernels`](https://github.com/huggingface/kernels/tree/main/skills/cuda-kernels) +- [CUDA Kernels Skill in `kernels`](https://github.com/huggingface/kernels/tree/main/kernel-builder/skills/cuda-kernels) - [HuggingFace Kernel Hub Blog](https://huggingface.co/blog/hello-hf-kernels) - [We Got Claude to Fine-Tune an Open Source LLM](https://huggingface.co/blog/hf-skills-training) - [We Got Claude to Teach Open Models](https://huggingface.co/blog/upskill) diff --git a/eee-community-evals.md b/eee-community-evals.md index 00c7d77d22..bcfecc4fbb 100644 --- a/eee-community-evals.md +++ b/eee-community-evals.md @@ -105,7 +105,7 @@ Hugging Face stores eval scores in the model repo as a YAML under `.eval_results Submit your full records to [the EEE datastore](https://huggingface.co/datasets/evaleval/EEE_datastore). -Utilizing EEE requires only one additional step, which the converter largely automates. The [community eval converter tool](https://github.com/evaleval/every_eval_ever/tree/main/tools/hf-community-evals) can be found in the GitHub repository. To process a collection, execute the following: +Utilizing EEE requires only one additional step, which the converter largely automates. The [community eval converter tool](https://github.com/evaleval/every_eval_ever/blob/main/every_eval_ever/tools/hf_community_evals.py) can be found in the GitHub repository. To process a collection, execute the following: ```shell uv run tools/hf-community-evals/community_evals_converter.py MMLU-Pro \ diff --git a/gemma4.md b/gemma4.md index 7f6c9d3a04..feaa27f8ef 100644 --- a/gemma4.md +++ b/gemma4.md @@ -633,7 +633,7 @@ mlx_vlm.generate \ --kv-quant-scheme turboquant ``` -For audio examples and more details, please check [the MLX collection](https://hf.co/mlx-community/gemma-4). +For audio examples and more details, please check [the MLX collection](https://huggingface.co/collections/mlx-community/gemma-4). ### Mistral.rs diff --git a/moe-transformers.md b/moe-transformers.md index 13ff1a0bab..68a58e5b3d 100644 --- a/moe-transformers.md +++ b/moe-transformers.md @@ -130,7 +130,7 @@ to: ### Dynamic Weight Loading with `WeightConverter` -The central abstraction introduced by this refactor is **dynamic weight loading** via a [`WeightConverter`](https://huggingface.co/docs/transformers/main/en/internal/weight_converter). +The central abstraction introduced by this refactor is **dynamic weight loading** via a [`WeightConverter`](https://huggingface.co/docs/transformers/main/en/weightconverter). `WeightConverter` lets us define: diff --git a/nvidia-reachy-mini.md b/nvidia-reachy-mini.md index 5a5691207a..13850e3810 100644 --- a/nvidia-reachy-mini.md +++ b/nvidia-reachy-mini.md @@ -38,7 +38,7 @@ We’ll be using the following: 1. A reasoning model: demo uses [NVIDIA Nemotron 3 Nano](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16) 2. A vision model: demo uses [NVIDIA Nemotron Nano 2 VL](https://huggingface.co/nvidia/NVIDIA-Nemotron-Nano-12B-v2-VL-BF16) 3. A text-to-speech model: demo uses [ElevenLabs](https://elevenlabs.io) -4. [Reachy Mini](https://www.pollen-robotics.com/reachy-mini/) (or [Reachy Mini Simulation](https://github.com/pollen-robotics/reachy_mini/blob/develop/docs/platforms/simulation/get_started.md)) +4. [Reachy Mini](https://www.pollen-robotics.com/reachy-mini/) (or [Reachy Mini Simulation](https://github.com/pollen-robotics/reachy_mini/blob/main/docs/source/platforms/simulation/get_started.md)) 5. Python v3.10+ environment, with [uv](https://docs.astral.sh/uv/) Feel free to adapt the recipe and make it your own \- you have many ways to integrate the models into your application: diff --git a/revamped-kernels.md b/revamped-kernels.md index 86ca7c0b9a..94ae266a6c 100644 --- a/revamped-kernels.md +++ b/revamped-kernels.md @@ -170,4 +170,4 @@ To solve this issue, kernels now link `libstdc++` dynamically. To ensure compati Our goal with the Kernels project is to serve both kernel developers and users of custom kernels. We’re always keen on receiving feedback from the community on how we can improve it. Don’t hesitate to contribute! -*Acknowledgements: Thanks to [Aritra](ariG23498) for reviewing the post.* \ No newline at end of file +*Acknowledgements: Thanks to [Aritra](https://huggingface.co/ariG23498) for reviewing the post.* \ No newline at end of file diff --git a/unsloth-jobs.md b/unsloth-jobs.md index 0935c9583b..4d06d9462c 100644 --- a/unsloth-jobs.md +++ b/unsloth-jobs.md @@ -96,14 +96,14 @@ Codex discovers skills through [`AGENTS.md`](https://developers.openai.com/codex Install individual skills with `$skill-installer`: ```text -$skill-installer install https://github.com/huggingface/skills/tree/main/skills/hugging-face-model-trainer +$skill-installer install https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer ``` For more details, see the [Codex Skills docs](https://developers.openai.com/codex/skills) and the [AGENTS.md guide](https://developers.openai.com/codex/guides/agents-md). ### Anything else -A generic install method is simply to clone the [skills repository](https://github.com/huggingface/skills) and copy the [skill](https://github.com/huggingface/skills/tree/main/skills/hugging-face-model-trainer) to your agent's skills directory. +A generic install method is simply to clone the [skills repository](https://github.com/huggingface/skills) and copy the [skill](https://github.com/huggingface/skills/tree/main/skills/huggingface-llm-trainer) to your agent's skills directory. ```text git clone https://github.com/huggingface/skills.git @@ -118,7 +118,7 @@ Once the skill is installed, ask your coding agent to train a model: Train LiquidAI/LFM2.5-1.2B-Instruct on mlabonne/FineTome-100k using Unsloth on HF Jobs ``` -The agent will generate a training script based on an [example in the skill](https://github.com/huggingface/skills/blob/main/skills/hugging-face-model-trainer/scripts/unsloth_sft_example.py), submit the training to HF Jobs, and provide a monitoring link via Trackio. +The agent will generate a training script based on an [example in the skill](https://github.com/huggingface/skills/blob/main/skills/huggingface-llm-trainer/scripts/unsloth_sft_example.py), submit the training to HF Jobs, and provide a monitoring link via Trackio. ## How It Works @@ -131,7 +131,7 @@ Training jobs run on [Hugging Face Jobs](https://huggingface.co/docs/huggingface ### Example Training Script -The skill generates scripts like this based on the example in the [skill](https://github.com/huggingface/skills/blob/main/skills/hugging-face-model-trainer/scripts/unsloth_sft_example.py). +The skill generates scripts like this based on the example in the [skill](https://github.com/huggingface/skills/blob/main/skills/huggingface-llm-trainer/scripts/unsloth_sft_example.py). ```python # /// script diff --git a/upskill.md b/upskill.md index 8caf3e99f5..734fa8a464 100644 --- a/upskill.md +++ b/upskill.md @@ -322,4 +322,4 @@ The approach works for any specialized task where you'd otherwise write detailed - [Upskill repo](https://github.com/huggingface/upskill) - [Agent Skills Specification](https://agentskills.io) -- [HuggingFace kernel-builder](https://github.com/huggingface/kernels/tree/main/builder) +- [HuggingFace kernel-builder](https://github.com/huggingface/kernels/tree/main/kernel-builder)