References¶
This page is the authoritative source list for the Concepts section. Each entry includes the author or organization, title, publication date where available, and URL. Concept pages cite these sources inline; when a source is referenced in multiple pages, the full citation lives here so it can be maintained in one place.
The list is grouped roughly by topic — architecture and design, cost and observability, security and incidents, and model selection — but is not strictly alphabetical. Use Ctrl+F / Cmd+F to find a citation by its short form (e.g., [Forbes, 2024]).
- Abstract Algorithms. "Build vs Buy: Deploying Your Own LLM vs Using ChatGPT, Gemini, and Claude APIs." Apr 2026. https://www.abstractalgorithms.dev/build-vs-buy-llm-self-host-vs-api
- Anthropic. "Building Effective AI Agents." Dec 2024. https://www.anthropic.com/engineering/building-effective-agents
- Anthropic. "Extended Thinking Guide." https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking
- AWS. "Multi-LLM routing strategies for generative AI applications." Apr 2025. https://aws.amazon.com/blogs/machine-learning/multi-llm-routing-strategies-for-generative-ai-applications-on-aws/
- Clarity. "How to Evaluate LLMs for Enterprise Use Without Getting Fooled by Benchmarks." Mar 2026. https://heyclarity.dev/blog/how-to-evaluate-llm-enterprise-use-without-getting-fooled-by-benchmarks/
- McLeod, Sam. "Patching NVIDIA's driver and vLLM to enable P2P on consumer GPUs." Feb 2026. https://smcleod.net/2026/02/patching-nvidias-driver-and-vllm-to-enable-p2p-on-consumer-gpus/
- Puget Systems. "Problems With RTX 4090 Multi-GPU and AMD vs Intel vs RTX 6000 Ada." https://www.pugetsystems.com/labs/hpc/problems-with-rtx4090-multigpu-and-amd-vs-intel-vs-rtx6000ada-or-rtx3090/
- Resilio Tech. "When to Self-Host AI Models vs. Use API Providers: A Decision Framework." Apr 2026. https://resiliotech.com/blog/when-to-self-host-ai-models-vs-use-api-providers-decision-framework
- TrueFoundry. "Intelligent LLM Routing: Cost & Quality-Aware Selection." Jun 2026. https://www.truefoundry.com/blog/llm-routing-cost-quality-aware-model-selection
- TrueFoundry. "LLM Benchmarking for Enterprise Production." May 2026. https://www.truefoundry.com/blog/llm-benchmarking-enterprise-production
- Anyscale. "LLM Router: Train and Deploy State-of-the-Art LLM Routers." https://github.com/anyscale/llm-router
- AWS. "MCP tool design: practical approaches and tradeoffs." Jul 2026. https://aws.amazon.com/blogs/machine-learning/mcp-tool-design-practical-approaches-and-tradeoffs/
- Boundev. "Long context vs RAG: when to use each in 2026." https://www.boundev.ai/blog/long-context-vs-rag-production-decision
- Gartner. "Over 40% of Agentic AI Projects Will Be Canceled by End of 2027." Jun 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- HumanLayer. "12-Factor Agents." https://github.com/humanlayer/12-factor-agents and https://www.hlyr.dev/blog/12-factor-agents
- Lin, X., et al. "REFRAG: Representation For RAG." arXiv, 2025. https://arxiv.org/pdf/2509.01092
- Liu, N. F., et al. "Lost in the Middle: How Language Models Use Long Contexts." Stanford / UC Berkeley / Samaya AI, 2023. https://arxiv.org/abs/2307.03172
- Manus. "Context Engineering for AI Agents: Lessons from Building Manus." https://manus.im/en/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus
- Milvus. "Claude Code Memory System Explained: 4 Layers, 5 Limits, and a Fix." 2026. https://milvus.io/blog/claude-code-memory-memsearch.md
- Nafiz, A. "How Claude Code Actually Remembers Things." https://ahammadnafiz.github.io/posts/How-Claude-Code-Actually-Remembers-Things/
- Ong, I., et al. "RouteLLM: Learning to Route LLMs with Preference Data." UC Berkeley / Anyscale / Canva. https://sky.cs.berkeley.edu/project/routellm/
- OpenAI. "Reasoning with the Responses API." https://platform.openai.com/docs/guides/reasoning
- Rausch, A. and Wittek, S. "Describing Agentic AI Systems with C4: Lessons from Industry Projects." arXiv, Mar 2026. https://arxiv.org/html/2603.15021
- Shopify Engineering. "Building Production-Ready Agentic Systems." Aug 2025. https://shopify.engineering/building-production-ready-agentic-systems
- Shou, C. (discovery) and community analyses. "Dive into Claude Code." VILA-Lab. https://github.com/VILA-Lab/Dive-into-Claude-Code
- Wang, S., et al. "Cognitive Architectures for Language Agents (CoALA)." arXiv, 2023. https://arxiv.org/abs/2309.02427
- Wu, Y., et al. "ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents." arXiv, 2026. https://arxiv.org/abs/2604.01664
- Zalesov, A. "A Practical Memo on Building LLM Agents." May 2026. https://medium.com/@zallesov/a-practical-memo-on-building-llm-agents-cc5179cb2e50
- Apollo / Daily Spark. "Cheaper Tokens, Bigger Bills." Jun 2026. https://www.apollo.com/wealth/the-daily-spark/cheaper-tokens-bigger-bills
- AWS. "Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region." Oct 2025. https://aws.amazon.com/message/101925/ (mirrored at https://gist.github.com/cosimo/53ee003ea00e4d6caa050f59d2a00a85)
- Hashimoto, M. "My AI Adoption Journey." Feb 2026. https://mitchellh.com/writing/my-ai-adoption-journey
- Böckeler, B. "Harness engineering for coding agent users." Martin Fowler, Apr 2026. https://martinfowler.com/articles/harness-engineering.html
- Business Insider. "Meta Employee Shares OpenClaw Email-Deletion Nightmare." Feb 2026. https://www.businessinsider.com/meta-ai-alignment-director-openclaw-email-deletion-2026-2
- DevOps.com. "When AI Goes Really, Really Wrong: How PocketOS Lost All Its Data." Apr 2026. https://devops.com/when-ai-goes-really-really-wrong-how-pocketos-lost-all-its-data/
- Fast Company. "Meta Superintelligence safety director lost control of her AI agent." Feb 2026. https://www.fastcompany.com/91497841/meta-superintelligence-lab-ai-safety-alignment-director-lost-control-of-agent-deleted-her-emails
- Forbes. "Klarna's AI Assistant Is Doing The Job Of 700 Workers." Mar 2024. https://www.forbes.com/sites/jackkelly/2024/03/04/klarnas-ai-assistant-is-doing-the-job-of-700-workers-company-says/
- Fortune. "As Klarna flips from AI-first to hiring people again." May 2025. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
- Fortune. "Tokens are getting cheaper, but AI costs keep climbing anyway." Jun 2026. https://fortune.com/2026/06/17/why-is-ai-spending-increasing-as-tokens-get-cheaper-jevons-paradox/
- Hardy, N. "The Confused Deputy." 1988. https://www.cs.utexas.edu/~witchel/380L/papers/hardy88confused.pdf
- IBM Research. "Unsupervised Cycle Detection in Agentic Applications." ICPE 2026. https://research.ibm.com/publications/unsupervised-cycle-detection-in-agentic-applications
- Ibrahim, M. "What Is Agent Observability?" Towards AI, Apr 2026. https://pub.towardsai.net/what-is-agent-observability-traces-loop-rate-tool-errors-and-cost-per-successful-task-dda2287f6c83
- Maxim. "7 Metrics You Should Track for AI Agent Observability." May 2026. https://www.getmaxim.ai/articles/7-metrics-you-should-track-for-ai-agent-observability/
- Menlo Ventures. "2025: The State of Generative AI in the Enterprise." Dec 2025. https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/
- OpenAI Engineering. "Harness engineering: leveraging Codex in an agent-first world." Engineering.fyi, Feb 2026. https://www.engineering.fyi/article/harness-engineering-leveraging-codex-in-an-agent-first-world
- OWASP. "LLM06:2025 Excessive Agency." https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/main/2_0_vulns/LLM06_ExcessiveAgency.md
- OWASP. "Top 10 for LLM Applications 2025." https://genai.owasp.org
- Piper Sandler / Stockwirex. "AI Token Costs Surge: Why Enterprise Bills Keep Climbing." Jul 2026. https://stockwirex.com/analysis/enterprise-ai-token-costs/
- Song, D., et al. "A Framework for Formalizing LLM Agent Security." OpenReview, 2026. https://openreview.net/forum?id=iQzd6qzIs5
- The Guardian. "Claude-powered AI agent’s confession after deleting a firm’s entire database." Apr 2026. https://www.theguardian.com/technology/2026/apr/29/claude-ai-deletes-firm-database
- WorkOS. "Delegated access for AI agents: The intersection rule explained." Jun 2026. https://workos.com/blog/delegated-access-ai-agents
- Grigorev, A. "How I Dropped Our Production Database and Now Pay 10% More for AWS." DataTalks.Club / AI Shipping Labs, Mar 2026. https://alexeyondata.substack.com/p/how-i-dropped-our-production-database
- Zhang, X., et al. "When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents." arXiv, 2026. https://arxiv.org/html/2607.01641
- ZenML. "Production Deployment of Toqan Data Analyst Agent: From Prototype to Production Scale." ZenML LLMOps Database, 2026. https://www.zenml.io/llmops-database/production-deployment-of-toqan-data-analyst-agent-from-prototype-to-production-scale
- Böckeler, B. "Harness Engineering - first thoughts." Martin Fowler, Feb 2026. https://martinfowler.com/articles/exploring-gen-ai/harness-engineering-memo.html
- OpenAI. "Harness engineering: leveraging Codex in an agent-first world." OpenAI, Feb 2026. https://openai.com/index/harness-engineering/
- Anthropic. "Effective Context Engineering for AI Agents." Sep 2025. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Hossain, M. S. "AI Agent Guardrails in Production: Input Filtering, PII Redaction & Prompt Injection Defense." Mar 2026. https://mdsanwarhossain.me/blog-ai-agent-guardrails.html
- Microsoft. "Agent Security with FIDES." Microsoft Learn, 2025. https://learn.microsoft.com/en-us/agent-framework/agents/security
- Redis. "Context Assembly: Building the Prompt the Model Sees." Redis blog, Jul 2026. https://redis.io/blog/context-assembly-building-the-prompt-the-model-sees/
- Claburn, T. "Cursor-Opus agent snuffs out startup's production database." The Register, Apr 2026. https://www.theregister.com/software/2026/04/27/cursor-opus-agent-snuffs-out-startups-production-database/5224442
- Pink, M., et al. "Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents." arXiv, 2025. https://arxiv.org/abs/2502.06975
- Luo, J., et al. "From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms." arXiv, 2026. https://arxiv.org/abs/2605.06716
- Zhang, D., et al. "Useful Memories Become Faulty When Continuously Updated by LLMs." arXiv, 2026. https://arxiv.org/abs/2605.12978
- Wolyra. "Enterprise LLM Selection Criteria: 2026 Framework." May 2026. https://wolyra.ai/enterprise-llm-selection-framework-2026/
- Hugging Face. "Security incident disclosure — July 2026." Jul 2026. https://huggingface.co/blog/security-incident-july-2026
- OWASP. "AI Agent Security Cheat Sheet." 2025. https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html
- ThousandEyes. "AWS Outage Analysis: October 20, 2025." Oct 2025. https://www.thousandeyes.com/blog/aws-outage-analysis-october-20-2025
- TrueFoundry. "LLM Failover & Load Balancing for Provider Outages." Jun 2026. https://www.truefoundry.com/blog/llm-failover-load-balancing-provider-outages
- Google. "LLM Agents vs. Workflows — and How Google ADK Gives You Both." Google Cloud, Apr 2026. https://medium.com/google-cloud/llm-agents-vs-workflows-and-how-google-adk-gives-you-both-7301d6fb1c4c
- Google. "Template agent workflows." Google ADK Docs, 2026. https://github.com/google/adk-docs/blob/main/docs/agents/workflow-agents/index.md
- OpenAI. "Orchestration and handoffs." OpenAI API docs, 2025. https://developers.openai.com/api/docs/guides/agents/orchestration
- CalibreOS. "LLM Router and Model Cascade: Cost-Aware Query Routing at Production Scale." 2026. https://www.calibreos.com/learn/genai-llm-router
- Chen, B., Zaharia, M., Zou, J. "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance." 2023. https://arxiv.org/abs/2305.05176
- Xue, R., et al. "R2-Router: A New Paradigm for LLM Routing with Reasoning." arXiv:2602.02823, 2026. https://arxiv.org/abs/2602.02823
- Li, Z., et al. "LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing." arXiv:2601.07206, 2026. https://arxiv.org/abs/2601.07206
- luyao618. "Claude Code Source Study — Memory Subsystem Overview." 2026. https://github.com/luyao618/Claude-Code-Source-Study/blob/main/docs-en/31-memory-subsystem-overview.md
- Jevons, W.S. "The Coal Question." 1865.
- Respan. "OpenAI vs Anthropic Pricing (2026): The Real Cost Math, Side by Side." May 2026. https://www.respan.ai/articles/openai-vs-anthropic-pricing-2026
- Stan, A.M. "AI token prices fell 98% but enterprise bills tripled." The Next Web, Jun 2026. https://thenextweb.com/news/token-prices-fell-98-enterprise-ai-bills-tripled-now-the-industry-wants-a-standards-body-to-explain-why
- AWS. "AGENTSEC04-BP02 Human-in-the-loop for critical decisions." AWS Well-Architected Agentic AI Lens, 2025. https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentsec04-bp02.html
- Solana Garden. "LLM Agent Loop Termination Explained: Stopping Criteria, Budget Caps and Production Guardrails." Jun 2026. https://solana.garden/guides/llm-agent-loop-termination-explained/
- Lin, X., et al. "REFRAG: Rethinking RAG based Decoding." arXiv, 2025. https://arxiv.org/abs/2509.01092
- Li, K., et al. "LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs." PMLR / ICML, 2025. https://proceedings.mlr.press/v267/li25dv.html
- Lumer, E., et al. "Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks." arXiv, 2026. https://arxiv.org/abs/2601.06007
- MLflow. "What Is Agent Observability? A 2026 Developer Guide." Jun 2026. https://mlflow.org/articles/what-is-agent-observability-a-2026-developer-guide/
- OpenTelemetry. "AI Agent Observability - Evolving Standards and Best Practices." Mar 2025. https://opentelemetry.io/blog/2025/ai-agent-observability/
- OpenTelemetry Semantic Conventions. "genai: define reasoning tokens attribute." Feb 2026. https://github.com/open-telemetry/semantic-conventions/pull/3383
- Uptrace. "OpenTelemetry for AI Systems: LLM and Agent Observability (2026)." Apr 2026. https://uptrace.dev/blog/opentelemetry-ai-systems
- Machine Learning Mastery. "AI Agent Tool Design: What Works and What Doesn't." Jun 2026. https://machinelearningmastery.com/ai-agent-tool-design-what-works-and-what-doesnt/
- Srinivasan, V. "Bridging Protocol and Production: Design Patterns for Deploying AI Agents with Model Context Protocol." arXiv, Mar 2026. https://arxiv.org/pdf/2603.13417
- TreeRouter. "MCP Server Production Guide: 8 Critical Pitfalls & Fixes." May 2026. https://api.treerouter.ai/en/blog/mcp-server-production-pitfalls-fixes-guide
- Song, Y. "Cache-Aware Prompt Compression: A Two-Tier Cost Model for LLM API Caching." arXiv, 2026. https://arxiv.org/abs/2607.15516
- Brown, Simon. "The C4 model for visualising software architecture." n.d. https://c4model.com/
- PlantUML. "PlantUML Language Reference Guide." Version 1.2025.0. 2025. https://pdf.plantuml.net/PlantUML_Language_Reference_Guide_en.pdf
- BPMN Quick Guide. "BPMN Modeling Best Practices." n.d. https://www.bpmnquickguide.com/quickguide/bpmn-quick-guide/bpmn-modeling-best-practices
- Lucidflow. "BPMN 2.0 Best Practices: The Twelve Mistakes Most First Diagrams Make, and How to Avoid Them." 2026. https://lucidflow.ai/blog/bpmn-best-practices
- Terrastruct. "D2: Declarative Diagramming." n.d. https://d2lang.com/
- DBML. "DBML Syntax — Core Database Markup." n.d. https://dbml.dbdiagram.io/docs
- Hill, Brenn. "Failure Recovery for Agent Loops: Retries, Rollback, and Resuming a Crashed Run." LoopRails, Jun 2026. https://looprails.dev/article-failure-recovery-agent-loops.html
- Nwaneri, Daniel. "How to Build a Production-Safe Agent Loop: From Exit Conditions to Audit Trails." freeCodeCamp, Jun 2026. https://www.freecodecamp.org/news/how-to-build-a-production-safe-agent-loop-from-exit-conditions-to-audit-trails/
- Cloud Security Alliance AI Safety Initiative. "Confused Deputy Attacks on Autonomous AI Agents." Mar 2026. https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-agent-confused-deputy-prompt-injection/
- Iusztin, Paul, and Louis-François Bouchard. "From 12 Agents to 1: AI Agent Architecture Decision Guide." Decoding AI, 26 Mar 2026. https://www.decodingai.com/p/from-12-agents-to-1-ai-agent-architecture-decision-guide
- Knowlee. "Single-Agent vs Multi-Agent: A Decision Framework (2026)." Knowlee Blog, 30 Apr 2026. https://www.knowlee.ai/blog/single-agent-vs-multi-agent-decision-framework
- Vectorize. "Claude Code Memory: Complete Guide to Persistence." 2026. https://vectorize.io/articles/claude-code-memory
- Novela. "Claude Code Deep Dive Part 6: How Claude Code Memory Is Designed." 2026. https://interwater.biz/blog/2026-04-06-claude-code-memory-architecture?lang=en
- Digital Applied. "Context Engineering: Agent Reliability Playbook 2026." Digital Applied blog, 2026. https://www.digitalapplied.com/blog/context-engineering-agent-reliability-playbook-2026
- Pan, T. "Token Budget as Architecture Constraint: Designing Agents That Work Under Hard Ceilings." tianpan.co, 2026. https://tianpan.co/blog/2026-04-13-token-budget-as-architecture-constraint
- Faros AI. "Harness Engineering: Making AI Coding Agents Work in 2026." Faros AI, May 2026. https://www.faros.ai/blog/harness-engineering
- DEV Community. "Harness Engineering: How to Build Production-Ready LLM Agents That Actually Work." DEV Community, May 2026. https://dev.to/monuminu/harness-engineering-how-to-build-production-ready-llm-agents-that-actually-work-20kc
- Zylos Research. "AI Agent Memory Architectures: From Context Windows to Persistent Knowledge." Zylos Research, 2026. https://zylos.ai/research/2026-04-05-ai-agent-memory-architectures-persistent-knowledge/
- JobsByCulture. "AI Agent Memory Systems: A 2026 Engineering Guide (Letta, LangMem, Mem0, Zep)." JobsByCulture, 2026. https://jobsbyculture.com/blog/ai-agent-memory-systems-guide-2026
- Bhairav, Suhas. "API-Based LLMs vs Self-Hosted LLMs: Production tradeoffs and deployment patterns." Jun 2026. https://suhasbhairav.com/blog/api-based-llms-vs-self-hosted-llms-fast-product-launch-vs-long-term-cost-control
- Morph Team. "Agent Tracing with OpenTelemetry (2026): How It Works and the Tools." Morph, Jun 2026. https://www.morphllm.com/agent-tracing
- OpenTelemetry. "Inside the LLM Call: GenAI Observability with OpenTelemetry." May 2026. https://opentelemetry.io/blog/2026/genai-observability/
- Singh, H. "Context Engineering Is Replacing Prompt Engineering: Building Production Context Pipelines for LLM Apps." DEV Community, Jul 2026. https://dev.to/hrsvd/context-engineering-is-replacing-prompt-engineering-building-production-context-pipelines-for-llm-14m8
- Clyro Content Team. "The $47K AI Agent Loop: A Complete Forensic Analysis." Apr 2026. https://clyro.dev/blog/the-47k-loop-a-complete-forensic-analysis/
- Orchesis. "I left my AI agent running overnight. Here's what I found in the morning." Mar 2026. https://orchesis.ai/blog/what-happens-ai-agent-runs-overnight
- Tianpan. "Long-Context vs RAG in 2026: Why It Is a Per-Feature Decision, Not an Architecture Religion." Tianpan Notes, 2026. https://tianpan.co/blog/2026-04-27-long-context-vs-rag-2026-decision-tree
- CallSphere. "Building Resilient AI Agents: Circuit Breakers, Retries, and Graceful Degradation." Mar 2026. https://callsphere.ai/blog/building-resilient-ai-agents-circuit-breakers-retries-graceful-degradation
- VeriSwarm. "Your LLM Provider Will Go Down. The Question Is Whether Your Agent Goes With It." May 2026. https://veriswarm.ai/blog/llm-provider-circuit-breakers
- Digital Applied. "LLM Model Routing in 2026: Cost-Quality Optimization." 2026. https://www.digitalapplied.com/blog/llm-model-routing-2026-cost-quality-optimization-engineering-guide
- LeanLM. "LLM Model Routing: Cheapest Capable Model Per Query." 2026. https://leanlm.ai/blog/llm-model-routing
- Tianpan. "The Inference Cost Paradox: Why Your AI Bill Goes Up as Models Get Cheaper." Apr 2026. https://tianpan.co/blog/2026-04-14-the-inference-cost-paradox
- Pontil. "MCP servers: a practical setup and architecture guide." May 2026. https://www.pontil.com/blog/mcp-servers-a-practical-setup-and-architecture-guide
- Kerkhoff, M. "MCP Integration Development Guide 2026." Context Studios, Jul 2026. https://www.contextstudios.ai/guides/mcp-integration-development-guide-2026