Data Engineer

Apollo is building out a data lake using an already-established medallion architecture, and expanding it as we bring on new clients and migrate existing ones. We're looking for a Data Engineer to lead this effort — owning the data lake's architecture and health, leading external data imports as we onboard new clients, and leading all API-based data imports, which flow directly into the data lake rather than our transactional database.

This is a leadership role on the data side: you'll be the person setting direction on how data flows in, how it's structured, and how ETL pipelines are built and maintained, while also staying hands-on with the pipelines and queries themselves.

Overview

Location: On-site (White Plains, NY)

Type: Full-time

Compensation: $80,000 - $90,000

Responsibilities
  • Lead the data lake build-out and expansion using our existing medallion architecture, onboarding both new and existing clients
  • Own external data imports as part of client onboarding
  • Lead all API-based data imports, ensuring API data is imported directly into the data lake rather than the transactional database
  • Design, build, and maintain ETL pipelines using Python
  • Write and optimize complex SQL queries against large datasets
  • Build web scraping tools to source external data as needed
  • Set technical direction and best practices for the team's data engineering work

Requirements

  • Experience with Azure Synapse (data lake)
  • Experience with Azure Data Lake Studio
  • Professional SQL experience, preferablyin a leadership role
  • Python experience, preferrably in production
  • Experience building and maintaining ETL pipelines
  • ETL experience using Python
  • Web scraping experience
  • Comfort owning deployments, not just writing code
  • Experience with Git/GitHub, Docker, or model deployment tools
  • Ability to clearly communicate insights to both technical and non-technical stakeholders
  • Ability to work independently and collaboratively in a fast-paced environment
  • Able to work 40 hours per week, in-office

Nice to Have

  • Experience with PySpark
  • Experience with Kafka
  • Experience with Flink
  • Microsoft 365 experience

Our Evaluation Method

    Our process includes a practical component — we're interested in how quickly you can get up to speed on an existing database/data lake architecture, and we'll ask you to work through some complex SQL queries.

    Please submit your resume and a cover letter to info@apollov2.com.

    ApolloV2 is an equal opportunity employer with a commitment to hiring people with diverse backgrounds. We do not discriminate based on age, civil or family status, disability, ethnicity, gender, race, religion, or sexual orientation.