AI for Data Engineers
AI is only as good as the data it retrieves. Build the pipelines that make AI systems reliable.
A hands-on track for data engineers who prepare data for AI systems. Retrieval-augmented generation and agents depend on well-processed, well-governed data: documents parsed correctly, content chunked and indexed sensibly, sensitive information controlled and everything kept fresh. This track covers document processing, embeddings and vector search, using language models inside pipelines, data quality and evaluation, governance and cost, with labs that include Arabic and bilingual content.
Curriculum
Data for AI: What Changes
How retrieval-augmented generation and agents use data, and why most AI quality problems start in the data layer. Structured versus unstructured sources and the new failure modes they create. The architecture of an AI data platform, from ingestion to retrieval.
Document Processing and Parsing
Extracting clean text and structure from PDFs, scans, tables, slides and web pages. OCR, layout handling and metadata capture. Handling Arabic and bilingual documents, including right-to-left text and mixed scripts.
Chunking and Embeddings
Chunking strategies and their trade-offs, choosing embedding models, including multilingual and open-weight options for data residency. Measuring how chunking and embedding choices change retrieval results.
Vector Search in Production
Vector indexes, hybrid search combining keywords and vectors, filtering and re-ranking. Running vector search in PostgreSQL with pgvector versus dedicated vector databases. Scaling, backups and operations.
Language Models Inside Pipelines
Using language models to extract fields, classify records and enrich data at scale, with structured outputs and validation. Handling errors, retries and costs when a model is one step in a batch pipeline.
Data Quality and Evaluation
Quality checks for AI-ready data: completeness, duplication, freshness and coverage. Building evaluation sets that connect data changes to answer quality. Monitoring data drift that silently degrades AI systems.
Governance, PII and Access
Lineage, access-aware retrieval, PII detection and masking, retention and deletion, and how data protection laws such as the UAE PDPL and GDPR apply to AI data. Making sure an AI system only retrieves what each user is allowed to see.
Capstone: An AI-Ready Data Platform
Design and build an end-to-end pipeline: ingestion, parsing, chunking, embedding, indexing, quality checks, governance and incremental refresh, with a cost estimate. Present your platform to a panel of HYVE engineers.
Learning Outcomes
- ✓Design data pipelines that prepare documents and records for retrieval and AI agents
- ✓Choose chunking, embedding and indexing strategies, and measure their effect on retrieval quality
- ✓Run and maintain vector search, in PostgreSQL with pgvector and in dedicated vector databases
- ✓Use language models inside pipelines to extract, classify and enrich unstructured data reliably
- ✓Build data quality checks and evaluation sets for AI-ready data
- ✓Apply governance to AI data, including lineage, access control, PII handling and retention
- ✓Keep AI data fresh with incremental updates and change detection
- ✓Control the cost and performance of AI data workloads at scale