Direkt zum Inhalt
Eine Person und eine Person, die auf Stühlen sitzt und auf einen Laptop schaut

Technology, Digital and Data

Senior Data Engineer

Ort Bangalore, Karnātaka / Chennai, Tamil Nādu, India
Datum der Veröffentlichung
Bewerben bis
Vertragsart Full time
Art der Tätigkeit Regular
Anforderungs-ID R0000364967

Beschreibung

Career Area:

Technology, Digital and Data

Job Description:

Your Work Shapes the World at Caterpillar Inc.

When you join Caterpillar, you're joining a global team who cares not just about the work we do – but also about each other. We are the makers, problem solvers, and future world builders who are creating stronger, more sustainable communities. We don't just talk about progress and innovation here – we make it happen, with our customers, where we work and live. Together, we are building a better world, so we can all enjoy living in it.

Position Overview

We are seeking a highly skilled and experienced Senior Data Engineer with strong expertise in PySpark, Azure Databricks, ETL/ELT, Microsoft Azure, and Microsoft Fabric. The candidate will design, develop, deploy, and maintain scalable enterprise data solutions, with a strong focus on Medallion Architecture, Lakehouse patterns, secure data sharing, Fabric capacity utilization, data quality, performance, governance, and operational reliability.

Key Responsibilities

  • Design, develop, and maintain scalable batch and streaming data pipelines using PySpark, Azure Databricks, and Microsoft Fabric.
  • Architect and implement Medallion Architecture across Bronze, Silver, and Gold layers for enterprise data processing and analytics.
  • Build robust ETL/ELT solutions for data ingestion, transformation, validation, reconciliation, and delivery across multiple source systems.
  • Develop and optimize PySpark and Spark SQL workloads for high-volume structured, semi-structured, and unstructured data.
  • Design and maintain Lakehouse and data lake solutions using Azure Data Lake Storage Gen2, Delta Lake, Microsoft Fabric OneLake, Fabric Lakehouse, and Warehouse.
  • Implement integration solutions using Azure Data Factory, Fabric Data Factory, data pipelines, notebooks, and Dataflows Gen2.
  • Design secure and governed data-sharing solutions across workspaces, domains, business units, and approved external consumers.
  • Implement reusable data products and cross-workspace sharing patterns using OneLake, OneLake shortcuts, Lakehouse, Warehouse, and semantic models.
  • Contribute to Microsoft Fabric capacity planning, workspace-to-capacity assignment, workload monitoring, utilization analysis, and performance optimization.
  • Monitor Fabric workloads using available capacity and workload metrics, identify resource contention, and recommend workload or scheduling improvements.
  • Design domain-aligned Fabric workspace structures with appropriate separation for development, testing, production, security, and ownership boundaries.
  • Implement data quality controls, monitoring, observability, lineage, error handling, reconciliation, and auditability across data pipelines.
  • Apply security best practices using managed identities, role-based access control, workspace roles, row-level or object-level controls where applicable, and secure secrets management.
  • Integrate data engineering solutions with Git-based source control and CI/CD pipelines for automated testing and deployment.
  • Optimize performance, scalability, reliability, and cost across Azure Databricks and Microsoft Fabric workloads.
  • Collaborate with Data Architects, Product Owners, Business Analysts, Data Scientists, QA engineers, governance teams, and platform teams in an Agile/Scrum environment.
  • Provide technical leadership, conduct design and code reviews, establish engineering standards, and mentor data engineers.

Required Skills & Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
  • 6+ years of proven Data Engineering experience, including delivery of enterprise-scale cloud data platforms.
  • Excellent hands-on expertise in PySpark, Spark SQL, DataFrame APIs, debugging, reusable framework development, and performance optimization.
  • Strong hands-on experience with Azure Databricks, including notebooks, jobs/workflows, clusters, Delta Lake, and production deployment patterns.
  • Practical implementation experience with Medallion Architecture, including Bronze, Silver, and Gold data layers.
  • Strong knowledge of ETL/ELT processes, ingestion patterns, incremental processing, transformation frameworks, and workflow orchestration.
  • Hands-on Microsoft Fabric experience, including Fabric Data Factory, notebooks, Lakehouse, Warehouse, OneLake, data pipelines, Dataflows Gen2, and semantic models.
  • Experience with Fabric capacity concepts, capacity assignment, utilization monitoring, workload analysis, performance tuning, and capacity-aware solution design.
  • Experience designing Fabric workspaces across domains and environments, including access, ownership, deployment, and workload-isolation considerations.
  • Strong experience implementing enterprise data-sharing patterns using OneLake shortcuts, shared Lakehouse or Warehouse data, governed data products, and semantic models.
  • Hands-on Azure experience, including Azure Data Lake Storage Gen2, Azure Data Factory, Azure Synapse Analytics, Azure Key Vault, and Azure Monitor.
  • Advanced SQL skills with experience in data modelling, data warehousing, query optimization, and performance tuning.
  • Experience with Delta Lake, Parquet, schema evolution, partitioning, and modern Lakehouse architecture patterns.
  • Strong understanding of data quality, metadata management, governance, lineage, security, privacy, and operational monitoring.
  • Experience with Azure DevOps or equivalent Git-based repositories and CI/CD pipelines for data engineering solutions.
  • Excellent problem-solving, troubleshooting, communication, and stakeholder-management skills.

Preferred Skills

  • Experience delivering Microsoft Fabric solutions in enterprise environments with multiple workspaces and shared capacity.
  • Experience interpreting capacity and workload metrics and recommending performance, scheduling, scaling, or workload-distribution improvements.
  • Knowledge of streaming architectures, event-driven processing, Eventstreams, and real-time analytics.
  • Azure Data Engineer, Azure Databricks, or Microsoft Fabric certifications are preferred.

Posting Dates:

September 14, 2026 - September 20, 2026

Caterpillar is an Equal Opportunity Employer. Qualified applicants of any age are encouraged to apply

Not ready to apply? Join our Talent Community.

LASST UNS DIE ARBEIT ANGEHEN

BLEIBEN SIE BEI DEN NEUESTEN JOBS UND CATERPILLAR-MELDUNGEN AUF DEM LAUFENDEN.

DER TALENTCOMMUNITY BEITRETEN
Eine Collage aus lächelnden Menschen