Data Engineering Intern

Apollo is building out a data lake using an already-established medallion architecture (Bronze / Silver / Gold), and expanding it as we bring on new clients and migrate existing ones. A large part of this work is moving business logic out of Tableau workbooks and into a centralized, documented Gold layer.

We're looking for a Data Engineering Intern to support that effort. You'll work alongside our data engineering team on pipeline development, SQL implementation, and — importantly — the discovery work that has to happen before anything gets migrated: figuring out how a report actually calculates a number today, writing an equivalent SQL definition, and documenting it so it can be validated and standardized.

This is a hands-on role. You won't be setting technical direction, but you will be doing real work that ships, with your output reviewed by senior engineers before it reaches clients.

Overview

Location: On-site/remote (White Plains, NY)

Type: Part-time (20 hours per week)

Compensation: $2500 monthly

Responsibilities
  • Reverse-engineer existing Tableau reports to identify how metrics are calculated — calculated fields, LOD expressions, filters, and context filters
  • Contribute to on external data imports as part of client management & onboarding
  • Write SQL implementations of those metrics and compare outputs against Tableau to confirm they match
  • Document metric definitions using our standard templates so report owners can confirm or correct proposed logic
  • Write and optimize complex SQL queries against large datasets; Write and test SQL queries against Gold layer views and tables
  • Build web scraping tools to source external data as needed
  • Support ETL pipeline development in Python under the direction of the data engineering team; Help maintain extraction and metadata tooling used across client environments

Requirements

  • Currently pursuing a degree in Computer Science, Data Science, Information Systems, Statistics, or a related field
  • Working SQL knowledge, including joins, aggregations, and window functions
  • Python experience, including familiarity with pandas
  • Comfort reading and reasoning about code or data structures you didn't write
  • Attention to detail — much of this work is confirming that two implementations produce identical results
  • Ability to communicate clearly in writing, including with non-technical stakeholders
  • Able to work 20 hours per week

Nice to Have

  • Exposure to Tableau or Power BI, whether as a builder or a user
  • Familiarity with cloud data platforms, ideally Azure Synapse or Azure Data Lake
  • Experience with Git / Github
  • Coursework or projects involving ETL or data pipelines
  • Microsoft 365 experience

Our Evaluation Method

    Our process includes a practical component — we're interested in how quickly you can get up to speed on an existing database/data lake architecture, and we'll ask you to work through some complex SQL queries.

    Please submit your resume and a cover letter to info@apollov2.com.

    ApolloV2 is an equal opportunity employer with a commitment to hiring people with diverse backgrounds. We do not discriminate based on age, civil or family status, disability, ethnicity, gender, race, religion, or sexual orientation.