Python
workflow / ML dataData preparation, service logic, automation, and reproducible analysis.
Concern: keeping transformations testable, reviewable, and repeatable outside a notebook.
Not supplied
A practical map of the technologies and responsibilities relevant to AI and data engineering work by Ketut Garjita. It is intentionally precise: a tool appears here with the concern it addresses, not as a decorative logo or an unsupported proficiency claim.
Select a technology to trace the engineering responsibility it represents. The matrix separates what a system must do—ingest reliably, evolve safely, schedule repeatably—from the name of the tool used to do it.
Interactive capability map
Use the filters to highlight related capability records and evidence slots.
Showing all technology records.
Data preparation, service logic, automation, and reproducible analysis.
Concern: keeping transformations testable, reviewable, and repeatable outside a notebook.
Not supplied
Distributed transformation of datasets that exceed a single-process workflow.
Concern: partitioning, shuffle cost, schema handling, and observable failure boundaries.
Not supplied
Representing recurring data work as dependency-aware workflows.
Concern: idempotent tasks, retries, backfills, scheduling semantics, and ownership.
Not supplied
Managed compute, storage, networking, identity, and deployment primitives.
Concern: provider-specific claims require naming the service, boundary, cost model, and operational context.
Provider not supplied
Bringing files, APIs, events, or operational records into a controlled pipeline.
Concern: incremental reads, rate limits, deduplication, late data, and source contracts.
Not supplied
Persisting raw, curated, and serving-ready data with an intentional access pattern.
Concern: partitioning, retention, schema evolution, access control, and recovery.
System not supplied
Making freshness, completeness, validity, and pipeline health visible.
Concern: actionable checks, thresholds, lineage, alert fatigue, and trustworthy ownership signals.
Not supplied
Building consistent datasets and features for training, evaluation, and inference.
Concern: leakage prevention, point-in-time correctness, reproducibility, and train/serve parity.
Not supplied
Tool names become useful when they reveal the decisions behind a system. These groups describe the boundaries a reviewer should look for in the project portfolio.
Reliable pipelines begin at the source boundary: incremental extraction, explicit schemas, retries, deduplication, and a clear response to malformed or late-arriving data.
Data should retain enough history to be explained and enough structure to be queried. Partitioning and transformation choices should follow access patterns, not habit.
Production confidence comes from more than a green task. Workflows need idempotency, backfill behavior, freshness expectations, meaningful checks, and alerts that help an operator decide what to do next.
AI systems inherit the weaknesses of their data layer. Useful evidence includes reproducible feature preparation, leakage controls, evaluation datasets, and a deliberate boundary between offline and online data.
Readable Python, versioned configuration, tests around transformations, documented assumptions, and reviewable changes are part of data engineering—not polish added afterwards.
The map below is a review checklist rather than an invented case study. When project records are available, each row should become a direct path from capability to implementation detail.
Select a technology above or use the links in each row. A strong project reference should answer what changed, why the design was chosen, how correctness was checked, and what operational trade-off remained.
Look for Python ingestion logic, source validation, incremental behavior, and a reproducible transformation boundary.
Case study not attachedLook for Spark workload shape, partition strategy, schema decisions, and evidence that performance was measured rather than assumed.
Case study not attachedLook for Airflow dependencies, retry and backfill behavior, freshness checks, and an operator-facing failure path.
Case study not attachedLook for the named provider service, identity and cost considerations, deployment boundary, and how serving data stays consistent with training data.
Provider and case study not attachedThese are useful review prompts, not hidden claims. The missing detail is itself important: it tells a hiring team what to ask for before treating a technology as demonstrated experience.
Bring a repository, architecture question, pipeline problem, or collaboration idea. A useful conversation can start with the constraints and the evidence—not just the tool list.