Asiman Data School

Data science and AI, from the very beginning.

A plain-language guide to how data, machine learning and AI work, and a six-month path to building them in Python.

Asiman Data School: Data and AI engineering education

Understanding Data Science and AI

A beginner's guide

Welcome! You have probably heard words like “AI”, “Machine Learning” or “ChatGPT” many times. This guide explains each one in simple words, with everyday examples and real uses in different industries.

The simple picture: think of these fields like a family

  • Data Science is the big family. It is all about learning useful things from data.
  • Machine Learning is a skill in the family. It teaches computers to learn from examples.
  • Deep Learning is a very strong type of Machine Learning that uses “neural networks”.
  • Computer Vision, NLP and LLMs are jobs that Deep Learning does: seeing, reading, and talking.
  • RAG is a trick that makes language models smarter and more trustworthy.

Data Science

Finding answers inside data

What is it?

Every day, the world creates a huge amount of data: shopping records, phone locations, hospital records, bank payments, weather measurements, and social media posts. By itself, data is just a big pile of numbers and words. Data Science is the work of turning that pile into useful answers.

An everyday example

Imagine you own a small shop and you write down everything you sell for a whole year. A data scientist looks at your notebook and tells you: “You sell the most ice cream on hot Fridays, your Tuesday sales are weak, and customers who buy bread usually also buy milk. So put milk next to the bread!”

How does it work?

A data scientist usually follows these steps:

  1. 1

    Ask a question.

    For example: “Why are we losing customers?”

  2. 2

    Collect data.

    Gather records from databases, websites, sensors, or surveys.

  3. 3

    Clean the data.Often the biggest part of the job

    Real data is messy. Some values are missing, some are typed wrongly, and some are repeated. Cleaning is often the biggest part of the job.

  4. 4

    Explore the data.

    Look at averages, patterns, and charts to understand what is going on.

  5. 5

    Analyze or build a model.

    Use statistics or Machine Learning to find deeper patterns or make predictions.

  6. 6

    Share the story.

    Explain the results with simple charts and clear words so that decision-makers can act.

Before cleaning

dateproductqtyproblem
4 JulIce cream12
4 JulIce cream12Repeated row
5 JulBread Missing value
5 JulMlik4Typed wrongly

After cleaning

dateproductqtywhat we did
4 JulIce cream12Copy removed
5 JulBread3Filled from the receipt
5 JulMilk4Spelling fixed
What cleaning looks like on a few rows of the shop's notebook (sample data).

What skills and tools are used?

  • Statistics (understanding numbers and chance)
  • Programming (mostly Python)
  • Databases (SQL)
  • Charts and dashboards

What it looks like in Python

Here is how a data scientist could ask the shop's notebook two of those questions with Python and the Pandas library. You will write code like this in Module 1.

shop_questions.py
import pandas as pd

# One row per item sold: receipt_id, date, weekday, product, quantity
sales = pd.read_csv("shop_sales.csv")

# Which weekday sells the most ice cream?
ice_cream = sales[sales["product"] == "ice cream"]
print(ice_cream.groupby("weekday")["quantity"].sum().sort_values(ascending=False))

# What do customers buy together with bread?
baskets = sales.groupby("receipt_id")["product"].apply(list)
with_bread = baskets[baskets.apply(lambda items: "bread" in items)]
print(with_bread.explode().value_counts().drop("bread").head(3))

Example code. The file shop_sales.csv stands in for the shop's year of sales.

Where is it used?

  • Business and Retail

    Finding best-selling products, deciding prices, understanding what customers want

  • Banking and Finance

    Spotting risky loans, detecting suspicious payments, analyzing markets

  • Healthcare

    Studying which treatments work best, planning hospital beds and staff

  • Sports

    Studying player performance, choosing game strategies

  • Government

    Planning city traffic, taxes, public services, and population needs

  • Marketing

    Measuring which advertisements work and who to show them to

  • Energy and Agriculture

    Predicting electricity demand, estimating crop harvests

Machine Learning

Teaching computers to learn from examples

What is it?

Normally, to make a computer do something, a programmer writes exact rules: “If this happens, do that.” But some problems are too complicated for rules. How would you write rules to recognize a spam email? Spammers keep changing their tricks.

Machine Learning (ML) solves this by letting the computer learn from examples instead of rules. You show it thousands of emails marked “spam” and “not spam”. The computer finds the patterns by itself, and soon it can judge new emails it has never seen.

Normal programming

Rules

Written by a programmer: “if the subject says WIN, it is spam”

Data

A new email arrives

Answer

Spam or not spam

Machine Learning

Data

Thousands of emails

Answers

Each one marked “spam” or “not spam”

Rules

Found by the computer itself: a model

The same spam problem, solved two ways.

Think of how a child learns what a “dog” is. Nobody gives the child a rulebook. They simply see many dogs, and soon they can recognize a new dog. Machine Learning works in a similar way.

The three main ways a machine learns

Supervised Learning (learning with an answer key). You give the computer examples together with the correct answers. For example: house details plus their real prices. Later, it can guess the price of a new house. It is like studying with a textbook that has answers at the back.

Predicting a number is called regression.

Price, temperature, sales.

Regression: predicting a house price from its sizeTwelve sample houses. A straight line fitted to them predicts about 171 thousand for a new 95 square meter house.60110160210260406080100120140Size (m²)Price (thousands)45 m², 82k52 m², 95k60 m², 110k68 m², 118k75 m², 140k83 m², 150k90 m², 162k98 m², 175k105 m², 190k112 m², 198k120 m², 215k130 m², 232kNew house: 95 m², predicted about 171kNew house: about 171k

Predicting a category is called classification.

Spam or not spam, sick or healthy.

Classification: spam or not spamSample emails placed by how many links they have and how many words like FREE or WIN they use. A dashed boundary separates spam from not spam.05100246810Links in the emailWords like FREE or WINNot spamSpam

Sample data for illustration. The line on the left is a real least-squares fit to the points shown.

Important ideas for beginners

Training and testing: We teach the model with one part of the data, then test it on a different part it has never seen. This checks that it truly learned and did not just memorize.

Training data: the model learns from itTest data: kept hidden until the end

Overfitting: When a model memorizes the training examples so closely that it fails on new ones. It is like a student who memorizes old exam answers but cannot solve new questions.

Learned the pattern

A simple line. It is close to the new points too.

Memorized the examples

Bent to touch every training point. It misses the new ones.

Training examplesNew examples it has never seenSize of the mistake
Sample data. Both models are computed from the same eight training points.

Features: The pieces of information given to the model, such as house size, location, and age.

Size85 m²
LocationCity center
Age12 years

Model

Prediction

House price

Where is it used?

  • Banking

    Deciding if a payment is fraud, scoring whether a person can repay a loan

  • E-commerce

    “You may also like...” product suggestions

  • Entertainment

    Movie, music, and video recommendations on streaming apps

  • Healthcare

    Predicting which patients are at high risk of a disease

  • Insurance

    Calculating fair prices and detecting false claims

  • Transport and Logistics

    Predicting delivery times and best routes, forecasting demand

  • Manufacturing

    Predicting when a machine will break before it actually does

  • Telecom

    Predicting which customers may leave and offering them deals

Deep Learning and Neural Networks

Machines that learn like a brain

What is it?

Deep Learning is an advanced type of Machine Learning. It uses a structure called a neural network, which is loosely inspired by how the human brain has billions of tiny cells (neurons) connected together.

How does a neural network work?

Imagine a long line of people passing a message, and each person slightly changes it before passing it on. A neural network is similar:

  • Input layer: Takes in the raw data (for example, the pixels of a photo).
  • Hidden layers: Many layers of small “neurons” in the middle. Each neuron looks at the information, gives it a certain importance (called a weight), and passes it forward. The word “deep” means there are many of these layers.
  • Output layer: Gives the final answer (for example, “this is a cat”).
catdogInputHidden layersOutput

Press a button to watch the network work.

Forward pass (the guess) Backpropagation (the correction)

How does it learn?

At first, the network guesses badly. After each guess, it checks how wrong it was. Then it goes backward through its layers and slightly adjusts the weights so that next time it is a little less wrong. This is called backpropagation. After millions of tiny corrections, the network becomes very good.

Why is it special?

In older Machine Learning, humans had to tell the computer what to look for. In Deep Learning, the network discovers what to look for on its own. When looking at faces, early layers may learn to see simple lines and edges, middle layers learn eyes and noses, and the last layers learn whole faces. This is why Deep Learning is so strong with images, sound, and text.

  1. Early layers

    learn to see simple lines and edges

  2. Middle layers

    learn eyes and noses

  3. Last layers

    learn whole faces

What does it need?

  • Lots of data
  • Powerful computers with special chips called GPUs

Common types of neural networks

CNN (Convolutional Neural Network)
Specialist for images.
RNN / LSTM
Older networks designed for sequences, such as sentences or time series.
Transformer
The newest and most powerful design. It is the engine behind modern chatbots and many image tools.

Where is it used?

  • Automotive

    The “brain” of self-driving car features

  • Healthcare

    Reading X-rays, MRI, and CT scans

  • Technology

    Voice assistants, photo apps, translation apps

  • Music and Media

    Creating music, voices, images, and videos

  • Science

    Predicting protein shapes, discovering new medicines, weather forecasting

  • Security

    Recognizing faces and detecting unusual activity

Computer Vision

Giving computers eyes

What is it?

For humans, looking at a photo is effortless. For a computer, a photo is just a giant table of numbers, where each number is the brightness or color of one tiny dot (pixel). Computer Vision teaches computers to understand what is in pictures and videos.

To a computer, this handwritten “0” is 64 numbers.

Each square is one pixel. 0 means blank paper and 16 means the darkest ink. A color photo works the same way, with three numbers per pixel: red, green and blue.

Real data: the first image in scikit-learn's handwritten digits dataset, which you will use in class.

What can computers do with images?

  • Image classification

    Say what is in the picture. (“This is a cat.”)

  • Object detection

    Say what is in the picture and where. (“Two cars and one person, here, here, and here.”)

  • Segmentation

    Color every pixel by what it belongs to: road, sky, tree, or building.

  • Face recognition

    Identify or verify a person from their face.

  • OCR (Optical Character Recognition)

    Read printed or handwritten text from a photo.

  • Pose tracking

    Follow body movements.

  • Image and video generation

    Create brand new pictures from a text description.

Where is it used?

  • Healthcare

    Finding tumors, fractures, or eye diseases in medical scans

  • Automotive

    Self-driving and driver-assist: detecting lanes, signs, pedestrians

  • Manufacturing

    Cameras on production lines that spot broken or faulty products

  • Agriculture

    Drones that check crop health, find weeds, and count fruits

  • Retail

    Checkout-free stores, shelf monitoring, virtual clothes try-on

  • Security

    Face unlock, airport checks, camera monitoring

  • Banking

    Reading cheques, verifying identity documents by selfie

  • Mapping and Environment

    Analyzing satellite images for forests, floods, and cities

  • Sports

    Tracking players and the ball, automatic replays

Natural Language Processing (NLP)

Teaching computers human language

What is it?

Humans communicate with language: speaking, writing, and reading. But language is tricky. Words have many meanings, people use slang, and sentences can be sarcastic. NLP is the field that helps computers understand, work with, and produce human language.

Key ideas in simple words

Tokenization: Breaking a sentence into small pieces (words or parts of words) so a computer can handle them.

Sentence

Tokenization helps computers read text.

Tokens

  1. Token
  2. ization
  3. ·helps
  4. ·computers
  5. ·read
  6. ·text
  7. .

Illustration. Every model has its own way of splitting text; the dot marks a space. Long or rare words are often split into parts, like “Token” + “ization”.

Embeddings: Turning words into numbers in a clever way so that words with similar meanings have similar numbers. “Happy” and “joyful” end up close together, while “happy” and “table” are far apart.

far apartclosehappyjoyfulgladcheerfultablechairdesksofa

Simplified picture. Real embeddings use hundreds of numbers per word, not two.

Context: The word “bank” means different things in “river bank” and “bank account”. Modern NLP looks at the surrounding words to understand the right meaning.

  • We sat on the river bank.

    means the land next to a river

  • I opened a new bank account.

    means a place that keeps money

What can NLP do?

  • Sentiment analysis

    Decide whether a review is positive or negative.

  • Translation

    Convert one language into another.

  • Summarization

    Turn a long article into a short summary.

  • Question answering

    Give answers to questions.

  • Speech recognition and voice assistants

    Turn speech into text and respond.

  • Named entity recognition

    Find names, places, dates, and amounts inside text.

  • Spam and toxic content detection

Where is it used?

  • Customer Service

    Chatbots that answer common questions 24/7

  • Healthcare

    Reading doctors' notes and medical papers to find key information

  • Legal

    Searching and summarizing thousands of contracts and cases

  • Finance

    Reading news and reports to understand market mood

  • Education

    Language learning apps, automatic essay feedback

  • Human Resources

    Sorting and matching CVs to job descriptions

  • Social Media

    Finding hate speech, understanding trends and opinions

  • Travel and Global Business

    Real-time translation

Large Language Models (LLMs)

The super-readers and super-writers

What are they?

An LLM is a very large neural network (a Transformer) that has read an enormous amount of text: books, websites, articles, and computer code. Its basic training task is surprisingly simple: guess the next word. For example, given “The sun rises in the...”, it learns to guess “east”. But when this is done billions of times on huge amounts of text, the model picks up grammar, facts, writing styles, reasoning patterns, and even programming skills.

The sun rises in the east

The model scores possible next words:

  • east82%
  • morning9%
  • sky5%
  • west1%

Illustrative scores. The model picks a likely word, adds it, and repeats, one token at a time.

ChatGPT, Claude, and Gemini are examples of tools built on LLMs.

Key ideas in simple words

Training (pre-training)
The long, expensive “reading” stage where the model learns from massive text.
Fine-tuning
Extra training to make the model polite, helpful, and good at following instructions.
Prompt
The message or instruction you give the model. A clear prompt usually gives a better answer.
Token
A small piece of text (a word or part of a word). Models read and write in tokens.
Parameters
The model's internal “knowledge dials”. Big models have billions of them.
Context window
How much text the model can keep in mind at one time, like its short-term memory.

What can LLMs do?

  • Write emails and stories
  • Explain difficult topics
  • Translate
  • Summarize long documents
  • Write and fix computer code
  • Answer questions
  • Brainstorm ideas
  • Act as a tutor
  • Control tools as “AI assistants” that complete tasks

Their weaknesses (important to know!)

  • Hallucination: Sometimes they say wrong things in a very confident way. Always double-check important facts.
  • Limited knowledge: They only know what they learned in training, so they may miss recent events or your company's private information.
  • No true understanding: They do not truly “understand” like a human. They are very good at patterns in language.

Where is it used?

  • Software Development

    Writing code, finding bugs, explaining code

  • Education

    Personal tutors, explaining lessons, creating practice questions

  • Marketing and Media

    Drafting articles, advertisements, social posts, scripts

  • Customer Support

    Smart assistants that handle real conversations

  • Healthcare

    Helping with paperwork, summarizing patient records (with human review)

  • Legal and Finance

    Drafting and reviewing documents, summarizing reports

  • Business Offices

    Meeting summaries, emails, data reports

  • Research

    Reading and comparing many scientific papers quickly

RAG (Retrieval-Augmented Generation)

Giving the AI an open book

What is it?

An LLM can only use what it learned during training. It does not know your company's private documents, and it does not know yesterday's news. RAG fixes this by letting the LLM look up information before it answers.

Think of two students taking an exam:

Closed book

Student A (plain LLM)

Has no books and answers from memory. Sometimes they guess wrongly.

“Refunds take 30 days, I think.” (A confident guess, no source.)
Open book

Student B (RAG)

Can open the right pages of a textbook, read them, and then answer. Their answers are more accurate and can point to the source.

“Refunds take 14 days.” Source: Refund policy, page 2.
Illustrative example answers.

How does it work? (4 simple steps)

  1. 1Prepare the library.

    Your documents (PDFs, web pages, manuals) are cut into small pieces. Each piece is turned into a list of numbers (an embedding) that captures its meaning. These are saved in a special vector database.

    Refund policy.pdf, cut into pieces[0.12, -0.48, 0.91, ...][-0.33, 0.27, 0.05, ...][0.64, 0.11, -0.72, ...]
  2. 2Search.

    When someone asks a question, the question is also turned into numbers. The system finds the document pieces with the closest meaning.

    “How long do refunds take?”[0.10, -0.51, 0.88, ...]Closest piece: the first one
  3. 3Add to the prompt.

    The best matching pieces are placed into the message sent to the LLM, together with the question.

    Question: How long do refunds take?
    Use this text: “Refunds are paid within 14 days...”
  4. 4Answer.

    The LLM reads those pieces and writes an answer based on them, often showing where the information came from.

    Refunds take 14 days.
    Source: Refund policy, page 2

The numbers are illustrative. Real embeddings have hundreds of numbers each.

Why is it useful?

  • Fewer wrong answers, because the AI answers from real documents.
  • Up-to-date and private knowledge without retraining the whole model.
  • Easy to update. Just add or replace documents.
  • Traceable. You can show the source, which builds trust.

Where is it used?

  • Companies (all sectors)

    Internal assistants that answer employee questions from company policies

  • Customer Support

    Bots that answer from the latest product manuals and help pages

  • Legal

    Searching through laws, contracts, and past cases

  • Healthcare

    Assistants that look up clinical guidelines and research

  • Education

    “Chat with your textbook” or course-notes assistants

  • Banking and Insurance

    Answering questions about complex rules, products, and regulations

  • Government

    Helping citizens find the right service or regulation

  • Software Teams

    Chatting with documentation and code libraries

How everything works together

One story. Imagine a modern online bank:

  1. A new customer signs up

    Computer Vision checks the selfie and ID card photo when a new customer signs up.

    Deep learning inside

  2. They make payments

    Machine Learning scores each payment and flags the ones that look like fraud.

  3. Months go by

    Data Science studies customer behavior and finds why some people close their accounts.

  4. A customer writes in

    NLP reads customer messages and understands whether the customer is angry or confused.

    Deep learning inside

  5. The bank replies

    LLMs write clear, friendly replies and chat with customers.

    Deep learning inside

  6. They ask about a rule

    RAG lets the chatbot look up the bank's latest rules and give correct answers.

Quick summary

Quick summary of the seven fields
Data ScienceFinding useful answers in dataExample: A shop learns which products sell best
Machine LearningComputers learn from examplesExample: Spam filter in your email
Deep LearningMachine learning with brain-like networksExample: Voice assistant understanding your speech
Computer VisionComputers understanding images and videoExample: Face unlock on your phone
NLPComputers understanding human languageExample: Translation apps
LLMsGiant language models that read and writeExample: ChatGPT or Claude
RAGLLMs that look up documents before answeringExample: A company chatbot that knows its own manuals

Questions beginners often ask

Do I need to know programming before I start?

No. Module 1 starts from the very beginning: variables, data types and operators, and the first lesson includes setting up Anaconda and Jupyter on your computer.

How much math do I need?

School math is enough to begin. The math that machine learning relies on (vectors, matrices, probability and statistics) is taught in Module 2, before machine learning starts.

Why Python?

Most data science and AI tools are built for Python, including the libraries used in this program: NumPy, Pandas, Matplotlib, Seaborn and scikit-learn. Python code is also short and readable, which makes it a good first language.

What computer do I need?

A regular laptop that can run Python is enough for most of the program, and the main tools are free. For training larger neural networks later on, free cloud notebooks such as Google Colab give access to GPUs, with usage limits.

How long is the program, and when are the lessons?

6 months: 52 lessons over 26 weeks, every Tuesday and Thursday, 19:00 to 21:00.

How is my progress checked?

Every Thursday, the second lesson of the week, there is an in-class Q&A and a Kahoot quiz on that week's material. Students who consistently do well can earn scholarships (a reduction in monthly program payments) and other prizes.