The VibeComputing Blog

27 articles · field notes on running AI in production

Agent Architecture · Infrastructure · Inference · Reliability

  1. Stop Prompting. Start Constraining.

    The agent's self-critique was flawless — and changed nothing. What fixed it: a regime gate, payoff arithmetic, and cooldowns. On why structure beats discipline…

  2. Blind, Not Biased: The Trading AI That Couldn't Read the Tape

    115 decisions in a falling market, every one of them "buy." The model wasn't misreading the trend — its inputs contained no direction at all. On missing features, falsification tests…

  3. The AI Agent That Graded Its Own Work an F

    Down 11% on its first session, the agent's own review opened with an F and the math behind it. On structural honesty: you don't get truthful self-assessment by asking harder…

  4. Your AI Agent Ignored You? It Wasn't Broken — It Was Outvoted

    An AI assistant stopped responding to its owner's direct messages. Logs showed delivery. The real cause: a background memory flush outvoted the human in a shared session. The fix i…

  5. The Fine-Tune That Forgot Its Own Name: Identity Bleed in Custom Models

    A fine-tuned model introduced itself as GPT-3.5 by OpenAI in a live demo. Fine-tunes don't inherit identity — base lineage bleeds through. The cure is training data, not prompt pat…

  6. The Fix That Fixed Nothing: Debugging the Wrong Layer

    A scheduled job kept skipping. The first fix was clean, logical, and completely irrelevant — the process never read it. How an unchanged error message points you to the layer you'r…

  7. Narrative Is Not Evidence: A Plausible Story Almost Cost 25 Hours

    An AI assistant gave confident 3D-print advice that would have wasted 2kg of filament and 25 machine-hours. The proof it was wrong was already in its own analysis output.

  8. Success-Status Laundering: When Agent 'OK' Means Nothing

    An AI agent returned success in 398 milliseconds and produced nothing. Three fleet failures where green status hid zero work — and the verification bar that fixes it.

  9. Debug From Where the Rejection Happens

    Two wrong theories, five hours. The rejecting server's log settled it in one line: clock skew. Why client-side debugging is theorizing in the dark.

  10. Estimates Dressed as Measurements

    The progress display sat at 35% for an hour. The print was fine. The number wasn't lying — it was never a measurement in the first place. How to tell derived metrics from ground tr…

  11. Shrink What's Active Before You Scale Out

    Same hardware, same quantization: a 35B sparse model on one node beat a dense model split across two — by 2-5x. The benchmark that inverts the scaling instinct.

  12. The Re-Launch Reflex: When Retrying Feels Like Progress

    Ten launch attempts. Every retry found a real bug. Every fix was correct. The plan was arithmetically dead from attempt #1 — and nobody ran the math. The failure mode that hides be…

  13. Silent Inflation: When Your Capacity Math Lies

    4-bit models that need 5x their size in memory. 20GB container images that take down 128GB servers. The hidden-expansion failure class that breaks every 'it fits' calculation — and…

  14. Quantization Is a Kernel Contract, Not a File Format

    A 21GB NVFP4 model that needed 100GB to run: why 'supported' quantization formats silently fail, how dequant fallbacks explode memory, and the three checks that prevent deployment …

  15. blog/inference-tuning-speculative-decoding.html

    Speculative decoding, continuous batching, and prefix caching — the serving-stack upgrades that outperform hardware purchases. Practical tuning guide for self-hosted LLM inference.

  16. blog/on-premise-ai-regulated-industries.html

    How healthcare, finance, and government teams run production AI on-premise — air-gapped LLM deployment, data residency compliance, and the real infrastructure math vs cloud APIs.

  17. Right-Sizing AI Infrastructure: Why Your Second GPU Might Be Slowing You Down

    A 295B model spread across two AI nodes ran 3x slower than a 26B model on one node. The active-parameters math every AI infrastructure buyer should know before scaling out.

  18. Self-Healing Infrastructure: From Auto-Restart to Autonomous Operations

    Self-healing infrastructure goes beyond auto-restart restarts. Learn how AI agents detect, diagnose, and remediate incidents autonomously — with guardrails that keep humans in cont…

  19. AI-Powered Infrastructure Observability: Beyond Traditional Monitoring

    How AI-powered infrastructure observability transforms incident response from reactive dashboards to proactive investigation. Real-time analysis, natural language queries, and auto…

  20. Vibe Slop is Coming. Here's How AI Infrastructure Avoids It

    The 'vibe slop' crisis is real — AI-generated code without governance creates technical debt, security holes, and outages. Here's how to use AI for infrastructure without the slop.

  21. DevSecOps in the Age of AI Agents

    How AI agents are transforming DevSecOps — from automated vulnerability detection to infrastructure hardening. Security best practices for AI-managed server fleets.

  22. blog/ai-infrastructure-cost-optimization.html

    AI infrastructure costs are exploding. Learn 7 proven strategies for AI infrastructure cost optimization — right-sizing, idle detection, spot instances, and how natural-language AI…

  23. blog/vibe-coding-security-risks.html

    38% of AI-generated code contains security flaws. Explore the real vibe coding security risks — and how zero-trust AI infrastructure tools solve them.

  24. blog/zero-trust-ai-agents.html

    A practical architecture for zero-trust AI agents in infrastructure: outbound-only agents, data obfuscation, approval gates, and BYOK. Learn how to safely give AI production access…

  25. AI Agent Security: How to Safely Give AI Access to Production Systems

    Giving AI agents access to production infrastructure is risky. This guide covers the 5 layers of AI agent security — from data obfuscation to approval gates to zero-trust agents.

  26. AI Infrastructure Automation in 2026: Beyond Scripts and Playbooks

    How AI infrastructure automation is replacing static playbooks with adaptive, natural language workflows. Learn about zero-trust agent architecture, data obfuscation, and approval …

  27. AIOps vs Traditional DevOps: Why the Gap Is Widening

    AIOps isn't DevOps with AI sprinkled on top. The gap between traditional DevOps and AI-native operations is widening — here's what changes when infrastructure speaks natural langua…

  28. Multi-Cloud Management Without Losing Your Mind

    Managing AWS, Azure, and on-prem servers doesn't need 5 dashboards. A practical guide to unified multi-cloud management with natural language and zero context switching.

  29. blog/natural-language-infrastructure.html

    The case for managing server fleets with natural language — why it works, where it fails, and what 'AI-augmented DevOps' actually means in practice.