
Senior Principal Data Engineer, Real-time Data (Remote)
Real job — pulled straight from Algoworks’s careers page · Verified September 7, 2026 · No reposts.
Job description
Algoworks is hiring a Senior Principal Data Engineer, Real-time Data (Remote) — a full-time, remote role. Apply directly on Algoworks's careers page below.
Senior Principal Data Engineer – Real-time Data
Department: PE-Cloud
Experience: 10+
- Lead the architecture and technical implementation of batch, full-load, incremental and real-time CDC pipelines.
- Design high-volume ingestion from SQL Server using CDC and Debezium.
- Build scalable event-driven pipelines using Azure Event Hubs and Databricks.
- Design and optimize Databricks pipelines for large-scale data ingestion and transformation.
- Implement robust error handling, retry, replay, checkpointing, recovery and idempotency.
- Design solutions for schema drift and schema evolution without disrupting downstream processing.
- Design and optimize Delta Lake / Delta Tables, including partitioning, compaction, data layout and performance optimization.
- Optimize pipeline throughput, latency, parallelism, resource utilization and processing windows.
- Establish monitoring and observability for CDC lag, connector health, consumer lag, pipeline failures, throughput and processing latency.
- Implement reconciliation and data-quality controls to ensure source-to-target completeness and accuracy.
- Provide technical direction, perform design/code reviews, mentor engineers and establish engineering best practices.
- Drive technical readiness for scaling ingestion across significantly more clients, databases, tables and data volumes.
- Bachelor’s or master's degree in computer science, Information Technology, Business, or related field (or equivalent practical experience).
- 10+ years of Data Engineering / Software Engineering experience.
- Strong hands-on experience with Databricks and Delta Lake.
- Strong experience designing and operating Databricks data pipelines at scale.
- Deep understanding of:
- Pipeline design and orchestration
- Error handling and recovery
- Schema drift
- Schema evolution
- Idempotent data processing
- Delta Tables
- Data partitioning and optimization
- Performance tuning
- Strong hands-on experience with SQL Server CDC, transaction logs, LSNs and high-volume transactional databases.
- Experience with Debezium SQL Server Connector, including configuration, offsets, snapshots, recovery and schema changes.
- Strong experience with Azure Event Hubs, including partitioning, consumer groups, scaling, throughput and checkpointing.
- Deep understanding of batch, micro-batch, streaming and event-driven data architectures.
- Strong experience with Python/PySpark, SQL, Azure Data Lake and distributed data processing.
- Experience designing production-grade solutions for retry, replay, fault tolerance, duplicate handling, reconciliation and observability.
- Strong performance engineering and troubleshooting skills across large-scale data pipelines.
- Ability to provide technical leadership, architecture guidance, mentoring and hands-on engineering support.
- 10+ years of Data Engineering / Software Engineering experience.
- 3+ years working with production-scale CDC or real-time streaming architectures.
- Strong production experience with Databricks and Delta Lake.
- Experience processing millions to billions of records.
- Experience with multi-client or multi-tenant ingestion architectures.
- Experience implementing Medallion / Bronze-Silver-Gold architectures.
- Experience with Apache Kafka / Kafka Connect and streaming ecosystems.
- Knowledge of Azure Data Factory, Azure Functions and Azure Monitor.
- Experience with Infrastructure as Code (Terraform/ARM/Bicep) and CI/CD for data platforms.
- Familiarity with Unity Catalog, Databricks Workflows and advanced Spark optimization.
- Establish a scalable architecture for batch and real-time ingestion.
- Scale pipelines across substantially more databases, clients, tables and data volumes.
- Improve Databricks pipeline performance and processing windows.
- Deliver reliable high-volume CDC without sustained lag, duplication, or data loss.
- Handle schema changes and schema drift without destabilizing ingestion.
- Ensure pipelines are idempotent and safely recoverable/replayable following failures.
- Optimize Delta Tables and downstream processing for performance and scalability.
- Provide clear technical leadership and mentoring for the ingestion engineering team.
Get Data Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Senior Data Engineer, Security & Governance
Frequently asked questions
Is Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks a remote job?
Yes, Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks is a remote position. This role is open to remote candidates.
What skills are required for Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks?
The required skills for Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks include: Databricks, Python, Spark, SQL, Kafka, Terraform, ARM, CI/CD.
What is the seniority level for Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks?
Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks is a Senior / Principal level position.
How do I apply for Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks?
You can view the full description and apply for Senior Principal Data Engineer, Real-time Data (Remote) at Algoworks on EchoJobs: https://echojobs.io/job/algoworks-senior-principal-data-engineer-real-time-data-5h1ov.