Spark & GenAI
Spark processes large datasets in bulk — a common way to generate embeddings or prepare training data at scale.
Prerequisites
Overview
Embedding a handful of documents can happen inline in an application. Embedding millions of them — or re-embedding a corpus after switching models — is a batch data-processing job, which is exactly what Spark is built for.
Where It Fits
Source Documents
Spark Job
Distributed processingEmbedding Model
Vector Database
Key Points
- Batch vs. inline embedding
- A large or changing corpus needs a scheduled batch job; a single new document can often be embedded inline when it’s created.
- Parallelism
- Spark distributes an embedding job across many workers, which matters once a corpus reaches millions of documents.
- Re-embedding
- Switching embedding models means re-processing the entire existing corpus — a job Spark is well suited to schedule and scale.
Interview Question
Why would you use Spark instead of just calling an embedding API in a loop?
A loop works fine for a handful of documents, but embedding millions of them — or re-embedding an entire corpus after switching models — needs distributed, fault-tolerant batch processing so a failure partway through doesn’t mean starting over, and so the work can be parallelized across many workers.
Explain It in 30 Seconds
Spark processes large document sets in parallel, batches, and with fault tolerance — well suited to generating or regenerating embeddings for a large corpus rather than embedding one document at a time inline.
Real-World Stack
Technologies commonly used to implement this in production.