Using Microsoft Web IQ to Enhance OpenWebUI: No-Code Web Search and Automatic Link Reading

In my previous article, I migrated the OpenWebUI database to Azure Database for PostgreSQL. In my day-to-day use of OpenWebUI, I kept running into two problems: No access to the latest information: Models have a knowledge cutoff. When I ask about recent releases, news, or prices, they either can’t answer or get it wrong. Pasted links are never opened: When I paste the URL of an article, a document, or a GitHub page into the chat and ask the model to summarize, translate, or compare it, the model doesn’t read the link at all....

September 28, 2026 · 12 min

Jev vs Laya: Comparing and Choosing Between Closed-Source and Open-Source System One Decision Models

Introduction In September 2026, a new kind of model that “doesn’t talk” suddenly became the talk of the AI world. On September 15, TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released Jev. It doesn’t chat, write articles, or generate code. It takes a piece of state and a set of predefined questions, and directly returns structured decisions with calibrated probabilities. Just three days later, on September 18, independent researcher Nandakishor M (Convai Innovations) open-sourced Laya....

September 23, 2026 · 25 min

Accelerating LLM Inference: Decoupling Prefill and Decode (PD Disaggregation)

As Large Language Models (LLMs) are widely deployed in dialogue systems, intelligent assistants, and Agent scenarios, the core challenge for inference systems has shifted from “can it run” to “how to run with lower latency, higher throughput, and greater stability”. Against this backdrop, PD Disaggregated (Prefill / Decode Decoupling) has gradually become a key architectural concept in large-scale online inference systems. This article will systematically explain what PD Disaggregated is, why it is needed, and the core advantages it brings to LLM inference systems from the perspective of model inference execution flow, without relying on any specific inference framework....

December 22, 2025 · 5 min

Optimizing Inference with Parameter/Data (P/D) Separation in vLLM Framework

Large language models often encounter GPU memory bottlenecks during inference deployment: Model parameters (P) can reach hundreds of GB and must remain resident in GPU memory. Input/output data (D) changes dynamically with each request but is often coupled with parameters on the same device, leading to imbalanced memory usage and limited scalability. To solve this problem, we can leverage the vLLM framework to implement Parameter/Data (P/D) Separation, improving the flexibility and throughput of inference systems....

September 29, 2025 · 5 min

Getting Started with Microsoft’s Latest Open-Source Long-Form Speech Model VibeVoice

What is VibeVoice? VibeVoice is a research framework released by Microsoft Research for long-form, multi-speaker, conversational speech synthesis. Target scenarios include entire podcast episodes, audio dramas, or interviews: it can maintain speaker consistency within a single generation and handle natural turn-taking. The model family includes multiple scales (e.g., 1.5B, 7B, etc.) and is available on Hugging Face as microsoft/VibeVoice-1.5B, along with model cards, weights, installation guides, and responsible use notes....

September 18, 2025 · 4 min