Senior AWS Data Engineer

AvenuecodeBrazilEngineerPublié le September 2, 2026

À propos de l'opportunité

We are looking for a Senior AWS Data Engineer to build and productionize data transformation layers for large-scale GA4 clickstream data. This role will focus on transforming raw data from an S3-based Bronze layer into scalable, domain-aligned Silver and Gold layers in Amazon Redshift, enabling analytics, Business Intelligence, and AI/RAG use cases.

This is a highly hands-on position focused on production delivery rather than prototypes. You will be responsible for building reliable data pipelines, designing analytics-ready data models, implementing data access controls, and optimizing Redshift workloads for downstream consumers such as Tableau, Power BI, and AI applications.

Responsabilités

  • Build and maintain production-grade data transformation pipelines using AWS Glue, Python, Matillion, or similar technologies.
  • Transform and flatten nested GA4/GA360 clickstream data, including complex struct and array fields, into queryable analytical tables.
  • Design and implement wide, denormalized, domain-aligned Silver-layer tables optimized for BI and analytical workloads.
  • Build Gold-layer aggregates on top of Silver data to support reporting, analytics, and AI/RAG use cases.
  • Implement column-level PII security and access controls using AWS Lake Formation, with distinct access models for BI, AI, and Data Science consumers.
  • Develop and deploy Redshift late-binding views as the primary consumption layer for Tableau and Power BI.
  • Ensure raw S3 data and underlying partitions remain protected from direct end-user access.
  • Optimize Amazon Redshift workloads, including distribution keys, sort keys, workload management, and query performance.
  • Reconcile and integrate new GA4 datasets with existing legacy clickstream data where appropriate.
  • Define and document data lineage, table grain, business definitions, and refresh cadence for production datasets.
  • Collaborate with technical and architectural stakeholders to make sound data architecture decisions.
  • Ensure clean documentation, knowledge transfer, and handoff at the end of the engagement.

Compétences requises

  • 5+ years of professional experience in Data Engineering, with recent hands-on production experience in AWS environments.
  • Strong experience with Amazon S3, AWS Glue, AWS Lake Formation, Amazon Redshift, and Redshift Spectrum.
  • Proven experience working with GA4 or GA360 data, particularly BigQuery-style nested and repeated structures.
  • Strong SQL and Python skills, with experience developing production-grade data transformations.
  • Hands-on experience with AWS Glue jobs, Matillion, or comparable data transformation frameworks.
  • Strong knowledge of Amazon Redshift performance optimization, including distribution keys, sort keys, late-binding views, and workload management.
  • Experience implementing column-level security and role-based access controls using AWS Lake Formation.
  • Ability to design data models that support BI, analytics, and AI/RAG workloads, beyond traditional normalized or Kimball-style modeling.
  • Strong understanding of data engineering best practices, including data quality, lineage, documentation, scalability, and production reliability.
  • Ability to work independently and collaborate effectively with architects, BI teams, data scientists, and other technical stakeholders.

Compétences souhaitées

  • Experience working with airline, travel, e-commerce, or high-volume customer clickstream data.
  • Experience with dbt or similar data transformation and modeling frameworks.
  • Familiarity with Tableau and Power BI, particularly their semantic-layer requirements and common issues caused by poorly designed schemas.
  • Experience preparing data for vector databases, LLM applications, or RAG pipelines.
  • Experience feeding AI or RAG applications from Redshift Gold-layer datasets.
  • Familiarity with modern data lakehouse and cloud data architecture patterns.
  • Experience working with large-scale event-based datasets and complex analytical workloads.