Law Wen Feng | Cloud and AI Architect
Search articles, topics, agents, cloud patterns… ⌘K
Subscribe

LLM Models

LLM Inference Optimization — Quantization, Speculative Decoding, and the 10x Cost Reduction Playbook

80% of AI GPU spend is now inference, not training. Here's the layered optimization playbook — quantization, KV cache compression, continuous batching, speculative decoding, and prompt caching — that delivers 10x cost reductions for Malaysian enterprises.

Aug 10, 2026 · 11 min read
Azure

Azure API Management as AI Gateway: Managing Access, Throttling, and Observability for LLM Endpoints

Architecture guidance for Azure API Management as AI Gateway, covering risks, governance decisions, and practical enterprise implementation considerations.

Jun 8, 2026 · 11 min read
LLM Models

LLM Deployment on Azure: Azure OpenAI vs. Self-Hosted Open-Source Models — A Cost and Architecture Decision Framework

A practical decision framework for Azure OpenAI versus self-hosted LLMs on Azure, including cost, security, control, and GPU operations.

May 8, 2026 · 9 min read
Page 1 of 1
© 2026 Law Wen Feng | Cloud and AI Architect · Having fun with Cloud and AI Agents.
Facebook LinkedIn X/Twitter RSS