<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Vllm-Sr on Wilson Wu</title><link>https://wilsonwu.me/en/tags/vllm-sr/</link><description>Recent content in Vllm-Sr on Wilson Wu</description><generator>Hugo -- 0.127.0</generator><language>en-US</language><lastBuildDate>Mon, 03 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://wilsonwu.me/en/tags/vllm-sr/index.xml" rel="self" type="application/rss+xml"/><item><title>Remote Embeddings in vLLM Semantic Router with Microsoft Foundry</title><link>https://wilsonwu.me/en/blog/2026/azure-foundry-embeddings-with-vllm-sr/</link><pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate><guid>https://wilsonwu.me/en/blog/2026/azure-foundry-embeddings-with-vllm-sr/</guid><description>1. Introduction In the design of vLLM Semantic Router, embedding is one of the core signals driving semantic routing.
The official tutorial has already demonstrated how to use Remote Embedding Providers to replace local embedding inference services:
👉 https://vllm-sr.ai/docs/tutorials/global/remote-embeddings/
However, in real production environments, a more critical question arises:
How can we use enterprise-grade, scalable embedding services to support semantic routing?
This article is based on the official tutorial scenario and combines Microsoft Foundry embedding models to demonstrate how to build a:</description></item></channel></rss>