Requesting team / customer: AMD FDE / Cohere Inference v-team (customer: Cohere — enterprise, on-prem/no-API; this benchmarking work supports Cohere's internal training pipeline evaluation and the longer-term goal of Cohere serving on AMD Instinct).
Jira Epic: pending — AIMODELS project key/access still being confirmed by requester; will link once created per Phase 1 of the intake process.
Model
HF repo: CohereLabs/c4ai-command-r7b-12-2024
Size / class: 7B dense, decoder-only
Framework: vLLM
Proposed MAD tier: vllm_extended (smoke-test tier — tp1, mc1), matching treatment of similarly-sized dense models (Llama-3.1-8B, Qwen3-8B) already in extended.yaml
Precision: auto / float16 on gfx942 (MI300 GEMM perf override, matching existing dense-model convention in the repo)
Hardware target
AMD Instinct MI300X (gfx942)
Workload / win-criteria
Standard MAD serving benchmark sweep (throughput/latency vs concurrency), used to validate Cohere's dense training checkpoint runs correctly and performantly on MI300X via vLLM.
Proposed changes
scripts/vllm/models.json — 1 new entry (pyt_vllm_command-r7b)
scripts/vllm/configs/extended.yaml — 1 new entry (command-r7b)
Draft config content available on request; no PR opened yet pending sign-off from requester and confirmation of Jira Epic link.
Open items before Phase 3 (Scoping) can close
Jira AIMODELS https://amd-hub.atlassian.net/browse/AIMODELS-1290
Filed per AMD's AI Models Intake Process (Confluence: MLSE space, "AI Models Intake Process"), Phase 2 "Technical Intake." This is 1 of 3 companion issues (command-r7b / command-a-plus / GLM-5.1-FP8) covering the same customer engagement.
Requesting team / customer: AMD FDE / Cohere Inference v-team (customer: Cohere — enterprise, on-prem/no-API; this benchmarking work supports Cohere's internal training pipeline evaluation and the longer-term goal of Cohere serving on AMD Instinct).
Jira Epic: pending — AIMODELS project key/access still being confirmed by requester; will link once created per Phase 1 of the intake process.
Model
HF repo: CohereLabs/c4ai-command-r7b-12-2024
Size / class: 7B dense, decoder-only
Framework: vLLM
Proposed MAD tier: vllm_extended (smoke-test tier — tp1, mc1), matching treatment of similarly-sized dense models (Llama-3.1-8B, Qwen3-8B) already in extended.yaml
Precision: auto / float16 on gfx942 (MI300 GEMM perf override, matching existing dense-model convention in the repo)
Hardware target
AMD Instinct MI300X (gfx942)
Workload / win-criteria
Standard MAD serving benchmark sweep (throughput/latency vs concurrency), used to validate Cohere's dense training checkpoint runs correctly and performantly on MI300X via vLLM.
Proposed changes
scripts/vllm/models.json — 1 new entry (pyt_vllm_command-r7b)
scripts/vllm/configs/extended.yaml — 1 new entry (command-r7b)
Draft config content available on request; no PR opened yet pending sign-off from requester and confirmation of Jira Epic link.
Open items before Phase 3 (Scoping) can close
Jira AIMODELS https://amd-hub.atlassian.net/browse/AIMODELS-1290
Filed per AMD's AI Models Intake Process (Confluence: MLSE space, "AI Models Intake Process"), Phase 2 "Technical Intake." This is 1 of 3 companion issues (command-r7b / command-a-plus / GLM-5.1-FP8) covering the same customer engagement.