Law Wen Feng | Cloud and AI Architect
Search articles, topics, agents, cloud patterns… ⌘K
Subscribe

LLM

Fine-Tuning Local LLMs in 2026 — Now Realistic on a Single Consumer GPU

Fine-tuning an 8B model on a single consumer GPU is now realistic. Here is the decision framework, hardware math, QLoRA workflow, and the pitfalls that sink real fine-tuning projects.

Aug 23, 2026 · 11 min read
LLM

Open-Source LLMs 2026: DeepSeek-V4, Kimi-K2.6, and the Enterprise Self-Hosting Decision

DeepSeek-V4 vs Kimi K2.6 for enterprise self-hosting: verified GPU math, real Azure Southeast Asia pricing, API break-even numbers, and a deployment guide for Azure-first teams.

Aug 20, 2026 · 10 min read
LLM

LLM Inference Optimization 2026: The Shift from Software Tweaks to Hardware-Software Co-Design

Why LLM inference optimization moved from software tweaks to hardware-software co-design, and what it means for Azure-first enterprise teams in 2026.

Aug 19, 2026 · 13 min read
Azure AI

Azure AI Foundry Model Catalog: Choosing From 10K+ Models for Enterprise Workloads

Microsoft's Foundry catalog lists 10,000+ models. Here's the decision framework I use to cut it down to the right one — with real pricing, deployment types, and a Malaysia West reality check.

Aug 16, 2026 · 8 min read
Azure

Azure API Management as AI Gateway: Managing Access, Throttling, and Observability for LLM Endpoints

Architecture guidance for Azure API Management as AI Gateway, covering risks, governance decisions, and practical enterprise implementation considerations.

Jun 8, 2026 · 11 min read
Page 1 of 1
© 2026 Law Wen Feng | Cloud and AI Architect · Having fun with Cloud and AI Agents.
Facebook LinkedIn X/Twitter RSS