Job Description
About the Role
The purpose of the Data Engineer I role is to support the design, development, testing, and documentation of data engineering solutions in line with LOA 2S expectations. The role collects, manages, and converts raw data into usable information for data scientists, analysts, and business stakeholders, while supporting data pipelines, data workflows, routine data tasks, and data quality controls that enable data-driven insights and operational efficiency.
What You Bring
Degree or Diploma in Computer Science, Engineering, Mathematics, or related field — NQF 5–6 (Essential) AWS Certification (Beneficial) 2–4 years relevant experience in a data team, including exposure to data science models and building/optimising data pipelines (Essential)
- Basic knowledge of data engineering concepts, data modelling, and databases
- Technical Skills such as:
- Python
- PySpark
- SQL
- dbt (Data Build Tool)
- Git / Version Control
- ETL pipeline development
- Data modelling
- Apache Airflow / Orchestration
- Data quality frameworks and observability
- Apache Iceberg / Delta Lake (open table formats)
- Terraform / Infrastructure as Code
- AWS Cloud — EMR, S3, Glue, Athena (Beneficial)
- Snowflake (Beneficial)
- Retail Operations experience
- Developing Specialist Capability — Supports the analysis, diagnosis, testing, resolution, and documentation of data solutions within an agile team
- Technical Aptitude — Strong passion and excitement for data, new technologies and solutions, and their range of possibilities and value for the business
- Self-Motivation — High level of drive to set, meet, and exceed goals and expectations. Uses own initiative in dealing with challenges
- Detail and Quality Focus — Has an affinity for structure and efficiency. Diligently watches over work processes, tasks, and outputs to ensure accuracy while promptly correcting quality concerns
- Communication — Communicates well both verbally and in writing. Able to simplify complex technical concepts for a variety of stakeholders
- Collaboration — Builds sound working relationships across the business. Able to work independently or collaboratively. Willing to coach/mentor others
- Resilience — Ability to work under pressure and tight time constraints, efficiently prioritising workloads in a high-volume, fast-moving environment
- Curiosity — Open to learning with a strong interest in data and discovery. Curious about exploring and answering business analytics questions
What You'll Do
- Creating data feeds from on-premises to cloud environments
- Supporting data feeds in production on a break-fix basis
- Building data warehouse layers and data products using dbt or similar transformation tools
- Manipulating data using Python and PySpark
- Processing data using distributed compute, particularly EMR Serverless Development for Big Data and Business Intelligence including automated testing and deployment
- Working with open table formats (Iceberg/Delta Lake)
- Writing and maintaining pipeline orchestration DAGs (Airflow)
- Using Git for version control, branching, pull requests, and code reviews
How well do you match?
Get an instant AI match score for this role — free, takes 3 minutes.
Tailor your CV for this role
The concierge rewrites your whole CV and writes a matching cover letter for this job — opens right here, nothing to paste.
Tailor My CV to This Job ✍️Free cover letter for this job
Upload your CV and get a tailored cover letter in seconds — free, no account needed.
Generate a Cover Letter 📝