Learn AI
Build your understanding from AI fundamentals to production AI systems.
Prefer a guided sequence? Explore Learning PathsData Engineering & GenAI
Move, process, and prepare the data that feeds AI pipelines.
7 conceptsKafka & GenAI
Kafka moves real-time events into AI pipelines — enrichment workflows, agent triggers, and downstream applications.
Intermediate · 4 minSpark & GenAI
Spark processes large datasets in bulk — a common way to generate embeddings or prepare training data at scale.
Intermediate · 4 minETL / ELT for AI
ETL/ELT pipelines extract, transform, and load the raw data that later gets chunked, embedded, or used for fine-tuning.
Intermediate · 4 minSnowflake & GenAI
Snowflake centralizes structured data that analytics and AI data-preparation pipelines both draw from.
Intermediate · 4 minDatabricks & GenAI
Databricks combines data engineering and ML workflows, often used to prepare and process data feeding model training or RAG ingestion.
Intermediate · 4 minPostgreSQL & GenAI
PostgreSQL stores application data and, via pgvector, can store embeddings in the same database rather than a separate system.
Intermediate · 4 minDocument Pipelines for AI
A document pipeline extracts text from source files (PDFs, wikis, tickets) before it ever reaches chunking or embedding.
Intermediate · 4 min