AI Pipeline Engineer

🏢 Bright Vision Technologies · all Bright Vision Technologies jobs
📍 United States
💰 USD 100,000 - 150,000 / annual
📅 Posted 2026-08-23 · via Himalayas
🏷 AI-Pipeline-Engineer,Data-Engineering,Machine-Learning-Infrastructure,Data-Pipeline-Engineering,AI-ML-Pipeline-Engineer,ML-Pipeline-Engineer,AI-Data-Pipeline-Engineering,ML-Data-Pipeline-Engineer,Generative-AI-Pipeline-Engineering,AI-Pipeline-Development,AI-ML-Engineer
Apply on original site ↗

AI Pipeline Engineer - Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: AI Pipeline Engineer
Location : 100% Remote (U.S.)
Position Type : Full-time, Direct W2
Salary Range : $100,000–$150,000 Annually
Experience Required : 6+ years

Sponsorship:  U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary
We are seeking an AI Pipeline Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.

Key Responsibilities
- Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement workflows.

- Build ingestion systems for diverse modalities including text, image, audio, video, and structured signals.

- Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale.

- Develop dataset versioning, lineage, and provenance tracking systems suitable for reproducible training.

- Build high-throughput data loading systems that maximize GPU utilization during training.

- Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement systems.

- Design storage architectures balancing cost, throughput, and latency across data tiers.

- Build evaluation dataset construction pipelines with strict integrity and contamination controls.

- Implement data privacy, redaction, and consent enforcement throughout the pipeline.

- Collaborate with ML researchers and engineers to align data systems with model development needs.

- Drive observability of data quality, drift, and pipeline health across the AI data estate.

- Optimize cost and performance through compression, format selection, and caching strategies.

- Document data systems, schemas, and operational procedures for broad internal use.

- Stay current with AI data infrastructure research and emerging open-source tools.

Required Qualifications
- Bachelor’s or Master’s degree in Computer Science or a related field.

- Six or more years of data engineering experience, with significant work supporting ML or AI workloads.

- Strong proficiency in Python and at least one JVM or systems language.

- Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.

- Hands-on experience operating petabyte-scale storage and pipeline systems.

- Strong understanding of distributed systems, data modeling, and storage formats.

- Experience with dataset versioning, lineage, and reproducibility for ML workflows.

- Familiarity with high-throughput data loading for accelerator-based training.

- Strong software engineering practices including testing, CI/CD, and code review.

- Excellent communication and cross-functional collaboration skills.

Preferred Qualifications
- Experience with multimodal datasets at large scale.

- Familiarity with data quality tooling and dataset evaluation methodology.

- Exposure to privacy-preserving data systems and regulated data handling.

- Open-source contributions to data infrastructure projects.

- Experience supporting frontier model training pipelines.

How to Apply
Would you like to know more about this opportun

← All remote jobs