Fine-tuning an 8B model on a single consumer GPU is now realistic. Here is the decision framework, hardware math, QLoRA workflow, and the pitfalls that sink real fine-tuning projects.
DeepSeek-V4 vs Kimi K2.6 for enterprise self-hosting: verified GPU math, real Azure Southeast Asia pricing, API break-even numbers, and a deployment guide for Azure-first teams.
Microsoft's Foundry catalog lists 10,000+ models. Here's the decision framework I use to cut it down to the right one — with real pricing, deployment types, and a Malaysia West reality check.
Architecture guidance for Azure API Management as AI Gateway, covering risks, governance decisions, and practical enterprise implementation considerations.