AI Software Masterclass Part 1: What Is AI Software? Types, Core Technologies & Working Principles Explained

AI Software Masterclass Part 1 - What Is AI Software Types Core Technologies and Working Principles
AI Software Masterclass Part 1: Understanding AI Software, Core Technologies & Working Mechanisms




Traditional software follows instructions. AI software makes predictions. That single difference explains almost everything else in this guide.

If you've ever wondered why ChatGPT doesn't behave like a calculator, or why an AI tool can be wrong with total confidence, the answer starts here. This is Part 1 of our five-part AI Software Masterclass, and it covers the foundation everything else in this series builds on.

By the end, you'll understand what actually happens between typing a prompt and receiving a response, without needing a computer science degree to follow along.

What Makes Software "AI Software"?

Every piece of software, at its core, takes an input and produces an output. What separates AI software from traditional software is how it gets there.

Traditional Software: Deterministic and Rule-Based

Traditional software runs on explicit instructions written by a programmer. If a condition is true, do one thing. If it's false, do another. This is called deterministic behavior — the same input always produces the exact same output, every single time.

A tax calculator is a good example. Enter the same income and deductions twice, and you get the identical result twice. There's no guessing involved.

AI Software: Probabilistic and Pattern-Based

AI software works differently. Instead of following hand-written rules, it learns patterns from large amounts of data, then uses those patterns to make a probabilistic prediction about the most likely correct output.

This is why the same prompt to an AI model can sometimes produce slightly different responses. It isn't following a fixed script — it's calculating the most probable answer based on everything it learned during training.
Aspect Traditional Software AI Software
Logic Explicit if/else rules Learned patterns from data
Output consistency Always identical for the same input Can vary slightly between runs
How it's built Programmed line by line Trained on large datasets
Failure mode Crashes or throws an error Can produce a confident but wrong answer
Neither approach is universally "better." A tax calculator should never guess. A tool writing marketing copy benefits from flexibility a rigid rule-based system could never offer.

The Building Blocks: Machine Learning, Deep Learning, and Neural Networks

These three terms get used almost interchangeably in casual conversation, but they describe a nested relationship, not three separate things.

Machine Learning: The Broad Category

Machine learning (ML) is the general approach of training a system to recognise patterns from data, rather than programming rules directly. A spam filter that learns which emails to flag, based on thousands of previous examples, is a classic machine learning application.

Deep Learning: A Specific Method Within ML

Deep learning is a subset of machine learning that uses structures called neural networks, arranged in multiple layers ("deep" refers to the number of layers). Each layer extracts increasingly abstract patterns from the data — early layers might detect simple shapes in an image, while deeper layers recognise entire objects.

Deep learning is what made the current wave of AI tools possible. It's particularly good at handling messy, unstructured data — text, images, audio — where writing explicit rules would be nearly impossible.

Neural Networks: The Underlying Structure

A neural network is loosely inspired by how neurons in a brain connect and pass signals to each other. In practice, it's a mathematical structure — layers of interconnected nodes, each adjusting numerical values called weights during training, until the network gets better at predicting the correct output.

The relationship, simply put: Machine learning is the umbrella category. Deep learning is one approach within it. Neural networks are the structure deep learning is built on.

Inside an LLM: How These Models Actually Work

Large Language Models (LLMs) — the technology behind tools like ChatGPT and Claude — deserve their own explanation, since this is where most of the current AI conversation is happening.

Step One: Breaking Text into Tokens

An LLM doesn't read words the way you do. It breaks input text into smaller units called tokens — sometimes a full word, sometimes just part of one. The word "unbelievable," for instance, might be split into a few smaller token pieces rather than treated as one unit.

Real world example of LLM Self Attention mechanism resolving context in ChatGPT
Figure 1: Practical demonstration of Transformer Self-Attention resolving pronoun context in LLMs.





Step Two: Converting Tokens into Embedding Vectors

Each token gets converted into a list of numbers called an embedding vector. This isn't arbitrary — the numbers are positioned so that tokens with related meanings end up mathematically close to each other. This is how a model captures the idea that "king" and "queen" are related concepts, without anyone explicitly telling it so.

Step Three: The Context Window

The context window is the maximum amount of text a model can consider at once, measured in tokens. A larger context window means a model can hold more of a conversation, document, or codebase "in mind" while generating a response, without losing track of earlier details.

Step Four: Transformers and Self-Attention

The architecture behind nearly every modern LLM is called a transformer. Its key innovation is a mechanism called self-attention, which lets the model weigh how relevant every other word in the input is to the word it's currently processing.

Real world example of Large Language Model Transformer Self-Attention mechanism in ChatGPT
Live demonstration showing how Transformer Self-Attention resolves context and pronoun reference in Large Language Models.





This is why an LLM can correctly figure out that in the sentence "the trophy didn't fit in the suitcase because it was too big," the word "it" refers to the trophy, not the suitcase. Self-attention lets the model connect that relationship across the sentence.

Step Five: Next-Token Prediction

At its core, generating a response is a repeated process: the model predicts the single most probable next token, adds it to the sequence, then predicts the next one after that, one token at a time, until the response is complete.

This explains a lot about how these tools behave. They aren't retrieving a pre-written answer from a database — they're constructing a response token by token, based on probability.

How These Models Are Trained

Building an LLM generally happens in stages:
  • Pretraining — the model learns general language patterns from massive amounts of text
  • Fine-tuning — the model is further trained on more specific, curated data to improve at particular tasks. Techniques like LoRA and QLoRA allow this fine-tuning to happen efficiently, adjusting a small portion of the model rather than retraining the entire thing from scratch
  • RLHF (Reinforcement Learning from Human Feedback) — human reviewers rate different model responses, and the model is further adjusted to favour the kinds of answers people rated as more helpful or accurate
Each stage shapes the model's behaviour. Pretraining builds raw capability. Fine-tuning and RLHF shape how that capability actually gets expressed in a conversation.

Why This Matters for How You Use These Tools

Understanding this pipeline explains a common source of confusion. A model can have enormous raw capability from pretraining, yet still respond in ways that feel overly cautious or oddly phrased — that's often the fine-tuning and RLHF stages shaping tone and behaviour, sometimes at the cost of directness.

It also explains why two AI tools built on similar underlying architecture can feel noticeably different to use. The base model might be broadly comparable, but the fine-tuning choices — what kinds of responses were rewarded during training — shape the personality and behaviour you actually experience.

Classifying Modern AI: Narrow, General, and Everything Between

Not all AI is the same kind of AI, and the terminology here gets used loosely in everyday conversation.

Narrow AI: What Almost Everything Today Actually Is

Narrow AI refers to systems built for a specific task. A recommendation engine, a fraud detection system, and a chatbot are all narrow AI, even the impressive ones. They perform well within their designed scope and don't generalise meaningfully beyond it.

General AI (AGI): Still a Roadmap, Not a Product

Artificial General Intelligence (AGI) describes a hypothetical system capable of understanding and performing any intellectual task a human can, across domains, without needing to be specifically trained for each one. As of now, AGI remains a research goal and an ongoing subject of debate, not something currently available in any commercial product. Timelines and predictions vary widely across the AI research community, and it's worth treating any confident date as speculation rather than an established fact.

Why the Narrow-vs-General Distinction Actually Matters

This isn't just a technical label. It shapes what you should reasonably expect from any AI tool you use. A narrow AI system trained for coding assistance can be genuinely excellent at that specific task while still making basic errors on something unrelated, like arithmetic or factual recall outside its training focus.

Understanding this helps set realistic expectations. An AI tool being impressively capable in one domain says nothing definitive about its reliability in a completely different one, because today's systems, however advanced, remain fundamentally narrow in what they were actually optimized to do well.

Generative AI vs Predictive/Analytical AI

This is one of the most practically useful distinctions for understanding the AI tools you actually encounter.
Type What It Does Examples
Generative AI Creates new content – text, images, audio, code ChatGPT, Midjourney, image and music generators
Predictive/ Analytical AI Analyzes existing data to forecast or classify Fraud detection, demand forecasting, spam filters
Generative AI produces something new that didn't exist before your prompt. Predictive AI looks at existing data and makes a judgment call — will this transaction likely be fraudulent, will demand for this product likely rise next month.

Many real-world AI systems actually combine both. A customer service tool might use predictive AI to route a ticket to the right department, then generative AI to draft the actual reply.

Bringing It Together: A Practical Example

Consider what happens when you ask an AI writing tool to draft an email.

The tool breaks your prompt into tokens, converts them into embedding vectors, and uses self-attention to understand how each part of your request relates to the others. It then predicts the most likely first token of a response, then the next, and the next, drawing on patterns learned during pretraining and refined through fine-tuning and human feedback.

None of this happens because the tool "understands" your request the way a person would. It happens because the underlying probability, shaped by enormous amounts of training data and careful adjustment, consistently produces text that reads as a coherent, helpful email.

Frequently Asked Questions

1. Is all AI software the same technology? No. AI software spans a wide range of approaches — from simple predictive models to massive language models. LLMs are one specific, currently prominent category, not a stand-in for AI as a whole.

2. Why does an AI model sometimes give a different answer to the same question?Because it's generating a probabilistic response, not retrieving a fixed answer. Small variations in the prediction process can lead to different, though often similarly reasonable, outputs.

3. Do I need to understand transformers and tokens to use AI tools effectively?          Not to use them day to day, but understanding these basics helps explain why AI tools behave the way they do — including their limits, like context window size and occasional confident mistakes.

What's Next in This Series

This part covered what AI software fundamentally is, and how the technology underneath it actually works. The next parts move from theory into practice.

Related Reading

Disclaimer: This article is for general educational purposes. AI technology and terminology continue to evolve, and specific tools or capabilities mentioned may change over time.


No comments:

Post a Comment

Popular Posts