Law Wen Feng | Cloud and AI Architect
Search articles, topics, agents, cloud patterns… ⌘K
Subscribe

LLM Models

Self-Hosting LLM Inference in 2026: vLLM vs SGLang vs llama.cpp Decision Guide

When self-hosted LLM inference beats API pricing in 2026 — with a decision matrix for vLLM, SGLang, and llama.cpp, working Azure deployment patterns, and the real GPU cost math for Malaysian enterprises.

Aug 26, 2026 · 10 min read
LLM

Open-Source LLMs 2026: DeepSeek-V4, Kimi-K2.6, and the Enterprise Self-Hosting Decision

DeepSeek-V4 vs Kimi K2.6 for enterprise self-hosting: verified GPU math, real Azure Southeast Asia pricing, API break-even numbers, and a deployment guide for Azure-first teams.

Aug 20, 2026 · 10 min read
Page 1 of 1
© 2026 Law Wen Feng | Cloud and AI Architect · Having fun with Cloud and AI Agents.
Facebook LinkedIn X/Twitter RSS