Databricks Data Engineer - India
Location: Remote
Work Hours: EST (Eastern Standard Time) aligned
Experience: 6โ9 years
Role Overview
We are looking for an experienced Databricks Data Engineer to support, maintain, and enhance existing Databricks-based data applications and pipelines. The role focuses on ensuring reliability, performance, and scalability of production Databricks workloads rather than building net-new platforms from scratch. You will work closely with data, analytics, and engineering teams to keep critical data applications stable, optimized, and aligned with business needs.
Key Responsibilities
-
Support and maintain existing Databricks applications, notebooks, jobs, and Delta Lake pipelines in production.
-
Monitor, troubleshoot, and resolve issues related to job failures, performance degradation, data quality, and cluster utilization.
-
Optimize existing Spark jobs, SQL queries, and Delta tables for cost, performance, and reliability.
-
Manage and improve Databricks workspace configurations, including clusters, job scheduling, access controls, and Unity Catalog (where applicable).
-
Implement and maintain data quality checks, logging, alerting, and basic observability for Databricks workloads.
-
Collaborate with stakeholders to understand requirements for enhancements or bug fixes on existing applications.
-
Perform incremental improvements, refactoring, and technical debt reduction on current Databricks solutions.
-
Ensure adherence to best practices around security, governance, and cost management within the Databricks environment.
-
Document existing pipelines, dependencies, and operational runbooks.
-
Participate in on-call or support rotations as needed to maintain production stability (within EST working hours).
Required Qualifications
-
6โ9 years of overall experience in data engineering, with strong hands-on experience in Databricks.
-
Solid proficiency in Apache Spark (PySpark and/or Scala) and SQL.
-
Proven experience supporting and optimizing production Databricks workloads (jobs, notebooks, Delta Lake, workflows).
-
Strong understanding of Delta Lake concepts (ACID transactions, time travel, optimization techniques such as Z-ordering, vacuum, optimize).
-
Experience with Databricks Job clusters, Interactive clusters, and performance tuning (partitioning, caching, shuffle optimization, autoscaling).
-
Familiarity with data modeling, ETL/ELT patterns, and production data pipeline support.
-
Experience working with cloud platforms (preferably Azure, AWS, or GCP) in the context of Databricks.
-
Ability to troubleshoot complex Spark and Databricks issues independently.
-
Strong communication skills and ability to work effectively in a remote, EST-aligned team.
Preferred Qualifications
-
Experience with Unity Catalog, Databricks SQL, or Lakehouse architecture.
-
Knowledge of CI/CD practices for Databricks (e.g., Databricks Asset Bundles, Git integration, Terraform/ARM templates).
-
Familiarity with orchestration tools (Airflow, Azure Data Factory, or Databricks Workflows).
-
Exposure to data quality frameworks, monitoring tools, or cost optimization initiatives on Databricks.
-
Experience supporting analytics or BI teams consuming Databricks data products.
Work Arrangement
- Fully remote
-
Must be available and productive during EST business hours
-
Collaborative remote environment with regular syncs and support responsibilities
Why Join UsWhy Join Us?
-
Join a team of industry veterans from Google, Meta, and top-tier tech companies.
-
Work on impactful, high-scale projects with leading global clients.
-
Enjoy a flexible, remote-first culture focused on innovation and excellence.
-
Competitive salary, equity options, and continuous learning opportunities.
-
Shape the future of cloud and AI infrastructure at a rapidly growing company.
Perks And Benefits Of Working With Us
- Internet allowance
- Laptop
- PF
- Paid PTO
- Annual Bonus
- Graduity
- Yearly