AI Software Masterclass Part 3: How AI Software Works — Architecture, LLM APIs, Vector Databases & RAG Setup

AI Software Masterclass Part 3 - How AI Software Works Architecture LLM APIs Vector Databases and RAG Setup
Figure: AI Software Masterclass Part 3 – Architecture, LLM APIs, Vector Databases & RAG Setup Overview





A chatbot demo and a production AI system look nothing alike underneath. The demo is one API call. The production system is a stack of six distinct layers, each one capable of quietly breaking the whole thing if it's built wrong.

Part 1 covered how AI models work internally. Part 2 covered which tools to actually use. This part covers what sits between them — the architecture that turns a capable model into a working business system.

If you've ever wondered why a company would need more than "just call the API," this is where that answer lives.

The 6-Layer Enterprise AI Stack

Every production AI system, regardless of industry, is built from the same basic layers stacked on top of each other.
Layer What It Handles
Infrastructure Servers, networking, storage, and security foundations
Compute GPUs or specialized chips that actually run model calculations
Model Layer The LLM itself, whether accessed via API or self-hosted
Vector Storage Databases that store and retrieve information the model can reference
Orchestration The logic connecting models, data, and tools into a working workflow
UI Layer The interface — chat window, dashboard, or app — where people interact with the system

Why Thinking in Layers Actually Helps

Most AI project failures trace back to treating this as one thing instead of six. A team might have an excellent model choice but a poorly designed vector storage layer, and the resulting system hallucinates constantly, not because the model is bad, but because it's being fed poor-quality retrieved information.

Diagnosing problems layer by layer, rather than blaming "the AI" broadly, is usually the fastest way to actually find what's wrong.

A Concrete Diagnostic Example

Consider a customer support assistant that keeps giving outdated pricing information. The instinctive reaction is often to blame the model for being unreliable. Walking through the layers usually reveals something different: the vector storage layer still holds embeddings generated from a pricing document that was updated weeks ago, and nothing re-indexed the new version.

The model isn't malfunctioning at all — it's accurately reporting exactly what it was given to work with. This kind of layer-by-layer thinking turns a vague "the AI is wrong" complaint into a specific, fixable problem.

Data Engineering for AI: Making RAG Actually Work

Retrieval-Augmented Generation, or RAG, is how most business AI systems answer questions about information the underlying model was never trained on — your company's documents, policies, or product data.

RAG Architecture and Vector Database Workflow Diagram - AI Bhaskar Guide
Figure: End-to-End RAG (Retrieval-Augmented Generation) Architecture Flow




How Chunking Actually Works

Before any document becomes searchable by an AI system, it needs to be broken into smaller pieces called chunks. This matters more than it sounds. A chunk that's too large dilutes the specific information the model needs. A chunk that's too small loses surrounding context that gives that information meaning.

Common approaches include splitting by fixed size, by natural document structure (paragraphs, sections), or by semantic meaning, where the text is divided based on where the topic actually shifts, rather than an arbitrary character count.

Why Chunk Overlap Matters

Most well-designed chunking strategies also include a small amount of overlap between consecutive chunks — repeating the last sentence or two of one chunk at the start of the next. Without this, a piece of information that spans exactly the boundary between two chunks can end up split in a way that neither chunk fully captures on its own.

This is a small technical detail with an outsized practical effect. Teams troubleshooting a RAG system that misses information "that's clearly in the document" often find the actual cause is a chunk boundary cutting the relevant sentence in half.

Vector Embeddings and Semantic Search

Once chunked, each piece of text gets converted into a vector embedding — the same kind of numerical representation covered in Part 1, positioned so that similar meanings sit close together mathematically. This is what allows a search for "refund policy" to also surface a chunk that says "return process," even without the exact word match.

Vector Database Comparison Chart Pinecone ChromaDB Qdrant - AI Bhaskar Guide
Figure: Key Vector Databases Comparison for AI Applications








Hybrid Search: Combining Two Approaches

Pure semantic search occasionally misses exact terms that matter — like a specific product code or legal clause number. Hybrid search combines semantic search with traditional keyword matching, catching both the conceptual match and the exact-term match in the same query.

Most production RAG systems today use hybrid search by default, rather than relying on semantic matching alone.

Choosing a Vector Database

Database Known For
Pinecone Fully managed, minimal setup overhead
ChromaDB Lightweight, popular for smaller projects and prototyping
Qdrant Strong performance at scale, open-source with a managed option
Milvus Built for very large-scale deployments, highly configurable

The right choice depends less on which is "best" and more on your team's existing infrastructure, expected data volume, and whether you want a fully managed service or more direct control.

Fine-Tuning vs. RAG: When to Actually Use Each

This is one of the most common points of confusion for teams starting out.
Approach Best For Trade-off
RAG Information that changes often, needs source attribution Adds retrieval latency, depends on data quality
Fine-tuning Teaching a consistent style, tone, or specialized behavior Doesn't update easily, requires retraining for new information
A simple rule of thumb: if the answer depends on information that might change next week — pricing, policies, current inventory — RAG is the better fit, since updating a database is far easier than retraining a model. If the goal is changing how a model responds, rather than what it knows, fine-tuning is more appropriate.

Many production systems use both together, rather than treating this as an either-or decision.

Closed-Source vs. Open-Source Models: The Real Trade-offs

This decision shapes nearly everything else about a system's architecture.

Commercial APIs

Using a commercial API — OpenAI, Anthropic, or similar — means calling a hosted model over the internet, without managing any of the underlying infrastructure yourself. This is the fastest path to a working system, with the trade-off that your data passes through an external provider's servers, and you're dependent on their pricing and availability.

Open Weights

Open-weight models — Llama, DeepSeek, Mistral, and others — can be downloaded and run on infrastructure you control. This gives you full data control and no per-token API costs, but shifts the burden of hosting, maintenance, and scaling onto your own team.

On-Premise and Sovereign AI Setups

For organizations with strict data residency requirements, running models entirely on owned or tightly controlled infrastructure — sometimes called sovereign AI — has become a genuine consideration, not just a theoretical one. This typically involves tools like vLLM or Ollama for serving the model, alongside dedicated GPU infrastructure sized to the model's actual parameter count.

The honest hardware reality: larger open-weight models require substantial GPU resources — often multiple high-end GPUs working together — which represents a real infrastructure investment, not a weekend project. Smaller open models can run on more modest hardware, making them a more realistic starting point for teams without existing GPU infrastructure.

A Reasonable Way to Approach This Decision

Rather than committing to a full self-hosted setup immediately, many teams start by prototyping with a commercial API to validate that the overall approach actually works for their use case. Only once that's confirmed does it make sense to evaluate whether the ongoing API costs at expected scale justify the upfront investment in self-hosted infrastructure. Committing to hardware before validating the underlying use case is one of the more expensive mistakes teams make in this space.

Orchestration Frameworks and Agent Protocols

Once you have a model and a data layer, something needs to coordinate them into an actual workflow. That's the job of orchestration frameworks.

The Major Frameworks

Framework Best Known For
LangChain General-purpose orchestration, wide integration ecosystem
LlamaIndex Strong RAG-focused workflows, document-centric applications
AutoGen Multi-agent conversation patterns
CrewAI Fast setup for role-based, collaborative agent teams
These frameworks aren't mutually exclusive competitors so much as tools suited to different starting points. CrewAI tends to get a basic multi-agent prototype running faster, while frameworks like LangChain offer more granular control once a system needs to scale or handle more complex state management.

The Model Context Protocol (MCP)

MCP has become a genuinely important piece of this landscape. Originally developed by Anthropic, it standardizes how AI models connect to external tools and data sources, solving what's sometimes called the "M×N problem" — without a shared standard, every model needs a custom integration for every tool, and that number grows unmanageably fast as both lists grow.

In December 2025, Anthropic donated MCP to the Linux Foundation, and it's now governed alongside a companion protocol through the Agentic AI Foundation. Every major framework mentioned above, along with tools like Cursor and Claude Code, now supports it as a default integration layer.

Model Context Protocol Workflow Diagram for Enterprise AI - AI Bhaskar Guide
Figure: Model Context Protocol (MCP) Workflow and System Integration




AI-to-AI Communication: A Separate Layer

Where MCP handles a model connecting to a tool, a related protocol called A2A (Agent-to-Agent), developed by Google DeepMind, handles communication between multiple AI agents working together. In a multi-agent system, individual agents typically use MCP to reach their own tools, while task handoffs between agents happen through A2A. The two are complementary layers, not competing standards.

Bringing the Stack Together: A Practical Example

Consider a company building an internal support assistant that answers employee questions using internal policy documents.

Documents get chunked and converted into embeddings, stored in a vector database like Qdrant. When an employee asks a question, hybrid search retrieves the most relevant chunks. An orchestration framework passes those chunks, along with the original question, to an LLM — either a commercial API or a self-hosted open model, depending on the company's data policies. The model generates a response grounded in the retrieved information, and MCP handles any additional tool calls, like checking a live HR system for a specific employee's leave balance.

Every layer in this example maps directly back to the six-layer stack described earlier. None of it works well in isolation.

Frequently Asked Questions

Do I need all six layers for a simple AI project?

No. A simple chatbot using a commercial API might only need the model, orchestration, and UI layers, skipping vector storage and heavy infrastructure work entirely. The full stack becomes relevant once a system needs to reference specific, changing information or run at meaningful scale.

Is RAG always better than fine-tuning?

Not always — they solve different problems. RAG suits frequently changing information; fine-tuning suits shaping consistent behavior or style. Many production systems use both rather than choosing one exclusively.

Why has MCP become so widely adopted so quickly?

Because it solves a genuinely painful integration problem that every team building AI agents ran into independently. Standardizing that connection once, rather than repeatedly, benefited every framework and lab enough that adoption spread quickly across the industry.

What's Next in This Series

This part covered how enterprise AI systems are actually architected, from raw infrastructure to agent communication protocols. The next part shifts to a different concern entirely.

Part 4 covers AI security risks — prompt injection, data poisoning, and the governance frameworks organizations use to manage them.

📖 Complete AI Software Masterclass Series:

Part 1: What Is AI Software? Types, Core Technologies & Working Principles Explained
Part 2: Top AI Software Tools for Coding, Content Creation & Design
Part 3: How AI Software Works: Architecture, LLM APIs, Vector Databases & RAG Setup (You are here)
Part 4: AI Software Security Risks: Data Privacy, Hallucinations & Enterprise Safety
Part 5: AI Software Pricing Models, Enterprise ROI & Future Industry Trends

Related Reading

Disclaimer: Technical tools, protocols, and platform capabilities reflect information available as of publishing and continue to evolve. Always confirm current documentation directly with each vendor or open-source project before implementation.

No comments:

Post a Comment

Popular Posts