#06
AI GLOSSARY/
42 terms — Fundamentals, architectures, techniques, models.
FUNDAMENTALS
7 termsKünstliche Intelligenz (KI / AI)
The umbrella term for systems that mimic human cognitive abilities like learning, problem-solving, and decision-making. AI isn't a single product — it's a field of methods, architectures, and applications.
Machine Learning (ML)
A subfield of AI where models learn from data — without being explicitly programmed. Instead of writing rules, you show the system examples and it finds patterns on its own.
Deep Learning
A subcategory of machine learning using neural networks with many layers. Enables recognition of complex patterns in images, speech, and text.
Neural Network (Neuronales Netz)
A computational model inspired by the human brain. Made up of layers of 'neurons' (nodes) that process signals and adjust weights to improve predictions.
Parameter
The learnable weights inside a model. GPT-4 has ~1.8 trillion parameters. More parameters = more capacity, but also more compute.
Training
The process by which a model learns on large datasets by minimizing errors. A model is exposed to data thousands of times until it makes accurate predictions.
Inference
Using a trained model to make predictions — like when you ask ChatGPT a question. Inference is significantly cheaper than training.
ARCHITECTURES
7 termsTransformer
The revolutionary architecture (2017, Google's 'Attention Is All You Need') that almost all modern LLMs are built on. Uses self-attention to understand relationships between words regardless of their position.
Self-Attention
A mechanism allowing the model to consider all other words in context when processing a single word. This is the core of the Transformer principle.
Large Language Model (LLM)
A Transformer-based language model with billions of parameters, trained on massive amounts of text. Examples: GPT-4, Claude, Gemini, LLaMA.
Diffusion Model
An architecture for generative image models (Stable Diffusion, DALL-E). Learns to iteratively remove noise from an image until a clear image emerges.
GAN (Generative Adversarial Network)
Two networks compete: a generator creates fakes, a discriminator tries to detect them. This competition produces realistic-looking images or videos.
Encoder / Decoder
Fundamental architecture principle: the encoder compresses inputs into an internal representation (embedding). The decoder generates output from it. Important for translations and summaries.
Mixture of Experts (MoE)
A scalable architecture where only a subset of parameters is activated for each input. GPT-4 and Gemini use MoE to be large without being proportionally expensive.
TECHNIQUES
10 termsPrompt Engineering
The art of crafting inputs (prompts) so an LLM produces optimal results. Includes techniques like chain-of-thought, few-shot, role prompting, etc.
Few-Shot Learning
Giving the model 2–5 examples within the prompt so it understands the pattern — without fine-tuning. Cheap and effective for structured tasks.
Chain-of-Thought (CoT)
Prompting the model to explain its reasoning step by step. Significantly improves accuracy on complex reasoning tasks.
RAG (Retrieval-Augmented Generation)
Combination of search and generation: relevant documents are retrieved live from a knowledge base and passed to the model as context. Enables current, source-based answers without retraining.
Fine-Tuning
A pretrained model is further trained on a smaller, specific dataset. Result: a model that behaves in an adapted style, domain, or behavior.
RLHF (Reinforcement Learning from Human Feedback)
Humans rate model outputs. These ratings train a 'reward model' that then refines the LLM through reinforcement learning. This is how ChatGPT & co. become helpful and safe.
Embeddings
Numerical vector representations of text, images, or other data in high-dimensional space. Similar concepts are close together. The basis for semantic search and RAG.
Vector Database
A database optimized for storing and searching embeddings (e.g. Pinecone, Weaviate, pgvector). Core component in RAG systems.
Context Window
The maximum amount of text (tokens) a model can process at once. GPT-4: 128k tokens. Claude: up to 200k tokens. Larger windows = more context, but more cost.
Token
The smallest unit an LLM processes — roughly ~¾ of an English word. 'Knowledgebase' = 2 tokens. Models think and compute in tokens, not words.
AI AGENTS & SYSTEMS
7 termsAI Agent
An LLM that doesn't just answer, but actively takes actions: searching, executing code, calling APIs, making decisions — in a feedback loop with the outside world.
Tool Use / Function Calling
An LLM can call defined functions (e.g. weather API, database query). The model output contains structured calls that a system then executes.
Agentic Loop
The pattern: think → act → observe → think. An agent iterates through this cycle until the task is done. The basis for autonomous systems like AutoGPT.
Multi-Agent System
Multiple specialized agents work together — one researches, one writes, one checks. Enables more complex tasks than a single agent.
System Prompt
An invisible instruction given to the model before every conversation. Defines the persona, behavior, constraints, and context of the assistant.
Grounding
Anchoring AI outputs to verifiable sources or real-time data. Reduces hallucinations and makes answers verifiable.
Halluzination
When an LLM confidently produces false facts. Happens because models predict probable token sequences — not look up facts. RAG and grounding help mitigate this.
MODELS & PROVIDERS
6 termsGPT-4o (OpenAI)
OpenAI's multimodal flagship model. Processes text, images, and audio. 'o' stands for 'omni'. The basis for ChatGPT Plus and the OpenAI API.
Claude (Anthropic)
LLM family from Anthropic, known for long context (200k tokens), precise reasoning, and Constitutional AI — an approach for safe AI behavior.
Gemini (Google DeepMind)
Google's multimodal LLM. Gemini 1.5 Pro supports up to 1 million token context. Integrated into Google Workspace and the Vertex AI platform.
LLaMA (Meta)
Meta's open-source LLM family. Can be run locally — no API key, no data sharing. The basis for many community models (Mistral, Vicuna, etc.).
Whisper (OpenAI)
OpenAI's open-source speech-to-text model. Transcribes audio in 50+ languages with high accuracy. Used in FINDELO for call transcripts, among other things.
Stable Diffusion
Open-source image generation model based on diffusion. Can be run locally. The basis for numerous commercial image tools and ComfyUI workflows.
INFRASTRUCTURE
5 termsGPU (Graphics Processing Unit)
Highly parallel processors optimized for matrix multiplications (the core of deep learning). NVIDIA H100s are the standard for LLM training. Extremely expensive.
API (Application Programming Interface)
An interface through which external systems communicate with a model. You send a request with your prompt — you get a response. The basis for all AI integrations.
Quantization
Reduces the precision of model weights (e.g. from 32-bit to 4-bit) to save memory and speed up inference — at the cost of slight quality loss.
LoRA (Low-Rank Adaptation)
An efficient fine-tuning method that modifies only a small fraction of parameters. Makes personalized fine-tuning possible on consumer hardware.
Latency
The time between sending a prompt and receiving the first response. Critical for real-time applications. Streaming reduces perceived latency.