Skip to content

Add Hugging Face and vLLM hosting options - #558

Merged
cheng-tan merged 1 commit into
mainfrom
shi-weili-docs-add-hf-vllm-hosting
Jul 22, 2026
Merged

Add Hugging Face and vLLM hosting options#558
cheng-tan merged 1 commit into
mainfrom
shi-weili-docs-add-hf-vllm-hosting

Conversation

@shi-weili

Copy link
Copy Markdown
Collaborator

Related Issue

N/A

Summary

Expand the model hosting guide with managed Hugging Face Inference Endpoints and self-hosted vLLM options alongside Microsoft Foundry Managed Compute.

Changes

  • Add Hugging Face deployment, authentication, model sizing, scale-to-zero, and cold-start guidance.
  • Add self-hosted vLLM commands and MagenticLite connection settings for Fara and MagenticBrain.
  • Reorganize Foundry Managed Compute as a third hosting option and document HTTP 503 cold-start troubleshooting.

How to Verify

  1. Review docs/model-hosting-guide.md and confirm each hosting option includes deployment and MagenticLite connection steps.
  2. Confirm the option and troubleshooting anchor links resolve to the corresponding sections.
  3. Run git diff --check origin/main...HEAD.

Checklist

  • Tests added or updated (if applicable)
  • Documentation updated (if needed)
  • Verified using the steps above

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>\nCopilot-Session: 031121eb-692f-4848-86a6-7803cb895333
@cheng-tan
cheng-tan merged commit 299fd20 into main Jul 22, 2026
13 checks passed
@cheng-tan
cheng-tan deleted the shi-weili-docs-add-hf-vllm-hosting branch July 22, 2026 16:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants