![]() |
Figure: AI Software Masterclass Part 3 – Architecture, LLM APIs, Vector Databases & RAG Setup Overview |
A chatbot demo and a production AI system look nothing alike underneath. The demo is one API call. The production system is a stack of six distinct layers, each one capable of quietly breaking the whole thing if it's built wrong.
Part 1 covered how AI models work internally. Part 2 covered which tools to actually use. This part covers what sits between them — the architecture that turns a capable model into a working business system.
If you've ever wondered why a company would need more than "just call the API," this is where that answer lives.
The 6-Layer Enterprise AI Stack
Every production AI system, regardless of industry, is built from the same basic layers stacked on top of each other.Why Thinking in Layers Actually Helps
Diagnosing problems layer by layer, rather than blaming "the AI" broadly, is usually the fastest way to actually find what's wrong.
A Concrete Diagnostic Example
The model isn't malfunctioning at all — it's accurately reporting exactly what it was given to work with. This kind of layer-by-layer thinking turns a vague "the AI is wrong" complaint into a specific, fixable problem.
Data Engineering for AI: Making RAG Actually Work
How Chunking Actually Works
Common approaches include splitting by fixed size, by natural document structure (paragraphs, sections), or by semantic meaning, where the text is divided based on where the topic actually shifts, rather than an arbitrary character count.
Why Chunk Overlap Matters
This is a small technical detail with an outsized practical effect. Teams troubleshooting a RAG system that misses information "that's clearly in the document" often find the actual cause is a chunk boundary cutting the relevant sentence in half.
Vector Embeddings and Semantic Search
Hybrid Search: Combining Two Approaches
Most production RAG systems today use hybrid search by default, rather than relying on semantic matching alone.
Choosing a Vector Database
The right choice depends less on which is "best" and more on your team's existing infrastructure, expected data volume, and whether you want a fully managed service or more direct control.
Fine-Tuning vs. RAG: When to Actually Use Each
A simple rule of thumb: if the answer depends on information that might change next week — pricing, policies, current inventory — RAG is the better fit, since updating a database is far easier than retraining a model. If the goal is changing how a model responds, rather than what it knows, fine-tuning is more appropriate.
Many production systems use both together, rather than treating this as an either-or decision.
Many production systems use both together, rather than treating this as an either-or decision.
Closed-Source vs. Open-Source Models: The Real Trade-offs
Commercial APIs
Open Weights
On-Premise and Sovereign AI Setups
The honest hardware reality: larger open-weight models require substantial GPU resources — often multiple high-end GPUs working together — which represents a real infrastructure investment, not a weekend project. Smaller open models can run on more modest hardware, making them a more realistic starting point for teams without existing GPU infrastructure.
A Reasonable Way to Approach This Decision
Orchestration Frameworks and Agent Protocols
The Major Frameworks
The Model Context Protocol (MCP)
In December 2025, Anthropic donated MCP to the Linux Foundation, and it's now governed alongside a companion protocol through the Agentic AI Foundation. Every major framework mentioned above, along with tools like Cursor and Claude Code, now supports it as a default integration layer.
AI-to-AI Communication: A Separate Layer
Bringing the Stack Together: A Practical Example
Documents get chunked and converted into embeddings, stored in a vector database like Qdrant. When an employee asks a question, hybrid search retrieves the most relevant chunks. An orchestration framework passes those chunks, along with the original question, to an LLM — either a commercial API or a self-hosted open model, depending on the company's data policies. The model generates a response grounded in the retrieved information, and MCP handles any additional tool calls, like checking a live HR system for a specific employee's leave balance.
Every layer in this example maps directly back to the six-layer stack described earlier. None of it works well in isolation.
Frequently Asked Questions
Do I need all six layers for a simple AI project?
Is RAG always better than fine-tuning?
Why has MCP become so widely adopted so quickly?
What's Next in This Series
Part 4 covers AI security risks — prompt injection, data poisoning, and the governance frameworks organizations use to manage them.
📖 Complete AI Software Masterclass Series:
Part 1: What Is AI Software? Types, Core Technologies & Working Principles ExplainedPart 2: Top AI Software Tools for Coding, Content Creation & Design
Part 3: How AI Software Works: Architecture, LLM APIs, Vector Databases & RAG Setup (You are here)
Part 4: AI Software Security Risks: Data Privacy, Hallucinations & Enterprise Safety
Part 5: AI Software Pricing Models, Enterprise ROI & Future Industry Trends
Related Reading
👉Moonshot Kimi K3: Open Weights, GitHub Setup & Databricks Integration 👉DeepSeek R1 vs OpenAI o1: The Open Model Debate 👉Claude Fable 5 for Multi-Agent Business Workflows 👉Claude Opus 4.8 Workflow & Monetization Guide
Disclaimer: Technical tools, protocols, and platform capabilities reflect information available as of publishing and continue to evolve. Always confirm current documentation directly with each vendor or open-source project before implementation.




No comments:
Post a Comment