Data Engineer
Overview:
Founded in 1898 and headquartered in Chicago, IL, GATX Corporation (NYSE: GATX) is an industry leader with 125+ years of success—success that is powered by our people. We are proud of our high-performance culture, hard-working and enthusiastic management team, and beautiful office space in the Willis Tower.
At GATX, we hire the best and offer our employees a dynamic, energetic, collaborative environment to enable them to make an impact from day one. Enjoy the perks and benefits of a global company with the close-knit culture and community of a much smaller one. In the same way we strive to empower our customers to propel the world forward, we are dedicated to providing our people with the tools and resources they need to advance in their careers.
This role will report to the Director, Analytics Engineering & Integration (DAEI) within IT’s “Data, Analytics, Reporting and Integration” (DARI) group and will help to evolve GATX’s data platform into a scalable, self-service data and AI ecosystem.
Although this is defined as a remote role and most work will be completed on a remote basis, the preferred candidate will live in the Chicago metropolitan area and can work from GATX’s Chicago office on a periodic basis.
Responsibilities:
Analytics Data Engineering
The Data Engineer will support new development requests, enhancements, and project initiatives that deliver data integration solutions aligned with business requirements. This role will collaborate with project teams, senior data engineers, architecture, infrastructure, analytics teams, business stakeholders, and external partners to support Data Mart, Data Warehouse, and cloud-based analytics platform needs in accordance with the organization’s enterprise data strategy.
This role will contribute to the quality and successful delivery of new and modified data components and will be expected to:
- Participate in all phases of the Software Development Life Cycle (SDLC) for data integration, ETL, and cloud-based data pipeline development, including requirements analysis, design, development, testing, deployment, and production support.
- Assist in designing, developing, testing, deploying, and supporting batch and near-real-time data pipelines on AWS and Databricks.
- Develop data transformations using SQL, Python, Spark, Delta Lake, and established data engineering patterns.
- Participate in requirements analysis, estimation, solution design, code reviews, testing, and production deployment.
- Develop and maintain Databricks notebooks, workflows, jobs, Delta tables, and related data components.
- Follow established architecture, coding, DevOps, security, governance, and documentation standards.
- Create unit tests and assist with integration, reconciliation, performance, and regression testing.
- Investigate data quality, pipeline, and production-processing issues, escalating complex problems when appropriate.
- Create and maintain source-to-target mappings, transformation rules, data flow diagrams, and operational documentation.
- Collaborate with senior engineers, architects, analytics teams, business stakeholders, infrastructure teams, and external partners.
- Build an understanding of supported business processes, source systems, reporting requirements, and upstream and downstream dependencies.
- Participate in platform upgrades, deployment automation, monitoring improvements, and operational process enhancements.
- Continuously develop expertise in AWS, Databricks, Spark, Python, SQL, Delta Lake, Unity Catalog, and modern data engineering practices.
Keeping the Lights On (KTLO)
- Assist in monitoring the health, performance, and reliability of data pipelines, ETL processes, and cloud-based data platform components.
- Support established monitoring, alerting, and operational procedures to help ensure data processing service levels and platform availability are maintained.
- Monitor daily production processes and work with IT administrators, support teams, and other DARI team members to investigate, troubleshoot, and resolve production issues in a timely manner.
- Maintain an understanding of established service level agreements (SLAs) and provide support that meets operational and business expectations.
- Participate in continuous improvement initiatives focused on reducing production incidents, improving system stability, and enhancing the overall customer experience.
- Diagnose and resolve data integration, ETL, and pipeline issues with guidance from senior team members, escalating complex issues as appropriate.
- Support the timely resolution of business-impacting data integration defects, failures, and processing issues to minimize disruptions to downstream reporting and analytics.
- Assist in root cause analysis activities and contribute to the implementation of corrective and preventive actions to improve platform reliability and operational efficiency.
- Create and maintain operational documentation, troubleshooting guides, and knowledge base articles to support ongoing platform operations and team knowledge sharing.
- Participate in a scheduled on-call rotation. Follow established troubleshooting and escalation procedures, working with senior engineers to resolve complex or business-critical incidents.
Innovation
- Stay informed about emerging data engineering technologies, cloud platform capabilities, and new releases of tools used by the Analytics and Data Integration teams, while developing expertise in AWS, Databricks, and related technologies.
- Assist in evaluating new features, enhancements, and upgrades to data integration and analytics platforms, providing input on potential improvements and implementation opportunities.
- Support the administration, configuration, testing, and maintenance of data integration tools and cloud-based data platform solutions, including both custom-developed and vendor-provided software.
- Participate in the implementation, testing, and rollout of platform upgrades, patches, and new capabilities to ensure system stability and business continuity.
- Apply and promote established development, data engineering, security, and operational best practices to support reliable, scalable, and maintainable solutions.
- Collaborate with senior engineers and team members to identify process improvements, automation opportunities, and operational efficiencies that enhance platform performance and delivery quality.
- Continuously expand technical knowledge and skills through training, self-development, and hands-on experience with modern data engineering tools, frameworks, and cloud technologies.
Qualifications:
- Bachelor’s degree in a quantitative or technical discipline such as Information Technology, Computer Science, Statistics, Economics, Mathematics, Engineering, or a related field.
- 1-3 years of experience in software development, data engineering, ETL development, data integration, analytics, or a related technical discipline within large, multi-platform enterprise environments (Windows, Unix, Oracle).
- Exposure to or experience with data integration, ETL processes, data warehousing, or cloud-based data platforms.
- Familiarity with Databricks, Apache Spark, Delta Lake, or similar big data technologies is preferred.
- Basic understanding of cloud platforms, preferably AWS, including services such as Amazon S3, IAM, and cloud-native data storage concepts.
- Experience working with SQL and relational databases for data extraction, transformation, and analysis.
- Basic programming experience in Python, SQL, or similar languages used for data engineering and analytics solutions.
- Understanding of data warehousing concepts, data modeling fundamentals, and data integration best practices.
- Familiarity with software development lifecycle (SDLC) processes and Agile development methodologies; experience with Jira or similar project management tools is a plus.
- Strong analytical, problem-solving, and troubleshooting skills with the ability to learn new technologies quickly.
- Ability to work effectively in a collaborative team environment and communicate technical concepts to both technical and non-technical stakeholders.
- Demonstrated attention to detail and commitment to delivering high-quality, reliable solutions.
Core Technical Skills
- Working knowledge of Apache Spark (PySpark preferred) and distributed data processing concepts.
- Experience developing data transformation and integration solutions using Python, SQL, or similar programming languages.
- Familiarity with relational databases and SQL development; experience with Oracle databases and PL/SQL is a plus.
- Understanding database concepts, data structures, data mapping, data transformations, and data modeling principles.
- Exposure to Databricks and cloud-based data engineering technologies, including:
- Delta Lake fundamentals
- Databricks Workflows and Jobs
- Data ingestion and transformation pipelines
- Experience or coursework involving batch data processing and data integration patterns, including:
- Incremental data loading
- Change Data Capture (CDC) concepts
- Streaming data fundamentals
- Architecture and Platform Experience
- Medallion Architecture (Bronze, Silver, Gold)
- Batch and streaming data processing concepts
- Data product and domain-oriented design concepts
- Dimensional data modeling
- Data quality and governance principles
- Data lifecycle management
- Understanding of modern data platform concepts, including data lakes, data warehouses, and lakehouse architecture.
- Familiarity with enterprise data architecture patterns such as:
- Basic understanding of:
- Contribute to scalable, maintainable, and reusable data solutions under the guidance of senior engineers
- DevOps and Automation
- Git-based version control systems
- CI/CD concepts and automated deployment practices
- Development, test, and production environment management
- Familiarity with software development best practices, including source control, code reviews, testing, and deployment processes.
- Experience or exposure to:
- Understanding of testing approaches for data pipelines and data quality validation.
- Governance, Security & FinOps
- Basic understanding of data security, access controls, and governance practices.
- Exposure to Databricks Unity Catalog or similar data governance tools is preferred.
- Understanding of cloud cost management, performance monitoring, and optimization concepts is a plus.
- Familiarity with data lineage and compliance requirements in enterprise environments.
- Familiarity with data visualization concepts and tools is preferred.
- Proficient with Microsoft Office Suite (Word, Excel, PowerPoint).
- Familiarity with enterprise data platform / lakehouse platform and transactional system integrations.
- Asset leasing industry experience is a plus, particularly within the Rail sector.
- Relevant certifications (e.g., Databricks Certified Data Engineer, AWS Certified Data Analytics, Azure Data Engineer) are preferred.
- Working knowledge of Unix shell scripting is a plus.
Other (i.e., physical requirements, travel, etc. that is not covered above):
- Occasional travel may be expected.
- Occasional work on nights and weekends to support production environment and meet project demands and deadlines.
Salary range:
As of the post date, the salary range for this position is:80800.00 USD - 96000.00 USD
This range is a reasonable estimate and takes into account several factors that are considered in making compensation decisions, including, but not limited to, geographic location, skill set, experience, education, training, internal equity, and other business needs.
This role may be eligible to participate in the Company’s short-term incentive plan, the details of which will be provided to the applicant upon hire.