Law Wen Feng | Cloud and AI Architect
Search articles, topics, agents, cloud patterns… ⌘K
Subscribe

LLM Models

Self-Hosting LLM Inference in 2026: vLLM vs SGLang vs llama.cpp Decision Guide

When self-hosted LLM inference beats API pricing in 2026 — with a decision matrix for vLLM, SGLang, and llama.cpp, working Azure deployment patterns, and the real GPU cost math for Malaysian enterprises.

Aug 26, 2026 · 10 min read
LLM Models

Claude Opus 4.8, Gemini 3.5 Flash GA, and the LLM Cost-Performance Frontier

The LLM market split into two clear tiers in mid-2026: frontier reasoning models and high-throughput workhorses. Here's how Malaysian enterprises should pick the right one for each workload.

Aug 11, 2026 · 8 min read
LLM Models

LLM Inference Optimization — Quantization, Speculative Decoding, and the 10x Cost Reduction Playbook

80% of AI GPU spend is now inference, not training. Here's the layered optimization playbook — quantization, KV cache compression, continuous batching, speculative decoding, and prompt caching — that delivers 10x cost reductions for Malaysian enterprises.

Aug 10, 2026 · 11 min read
LLM Models

LLM Inference in Production 2026: Quantization, KV Cache, Speculative Decoding, and the vLLM vs SGLang Decision

Running LLMs in production is a serious engineering discipline. The gap between naive and optimized inference is the difference between a viable product and an expensive failure.

Aug 2, 2026 · 6 min read
LLM Models

LLM Deployment on Azure: Azure OpenAI vs. Self-Hosted Open-Source Models — A Cost and Architecture Decision Framework

A practical decision framework for Azure OpenAI versus self-hosted LLMs on Azure, including cost, security, control, and GPU operations.

May 8, 2026 · 9 min read
Page 1 of 1
© 2026 Law Wen Feng | Cloud and AI Architect · Having fun with Cloud and AI Agents.
Facebook LinkedIn X/Twitter RSS