Senior Data Scientist, AI Scoring & Evaluation

🏢 Workera AI · all Workera AI jobs
📍 Remote · United States
📅 Posted 2026-08-27 · via RemoteIO
🏷 Data Science,Machine Learning,Artificial Intelligence,Python,SQL
Apply on original site ↗

Senior Data Scientist - Assessment Scoring & Evaluation
Everyone's racing to build AI. Workera exists for the 8 billion people who work alongside it.
While the world's attention is on creating new tools, someone has to solve the other side of the equation: the humans. The workforce is going through the biggest transformation in a generation and most organizations are navigating it blind, without the data to understand what their people can actually do, where the gaps are, or how to close them fast enough.
As Workera moves into higher-stakes decisions, you're the person who can prove a measurement is accurate, fair, and defensible. You'll own the measurement system that turns signals into the skill data our customers act on: the evaluation harnesses, the calibration methods, the quality monitoring behind every score. You'll also shape how our AI software itself is designed, from how components are orchestrated and prompted to the evaluation frameworks that decide whether they're good enough to ship.
You're not just running analyses, you're building the trust layer underneath a product that increasingly makes high-stakes calls about people's careers. If you want a role where your judgment becomes the thing customers, auditors, and your own engineering team rely on, and where the measurement problems are still being invented rather than optimized, this is it.
Workera's skills intelligence platform is critical infrastructure for the AI era: the layer that lets organizations understand, mobilize, manage, and develop their talent with precision. We're trusted by the Fortune 500, powered by proprietary AI agents, and built by a small, senior team, which means what you ship here has outsized reach.
WHY THIS ROLE EXISTS
Workera is scaling fast: more customers, more use cases, more fields evaluated, higher stakes on every measurement. That scale changes what quality means. What we once crafted and inspected by hand now needs monitoring that catches issues before customers do, and remediation that resolves them fast.
This role sits at the intersection of Engineering, Assessment Science, and Design, partnering closely with the product engineers who build our scoring pipeline. You're the data layer that pieces those disciplines together: the foundation the rest of our scoring is built on, and the reason a construct definition from Assessment Science actually turns into a number a customer can trust.
YOUR TEAM
You report to Dr. Taylor Sullivan, Workera’s VP of Product and Assessments, and work day to day with two groups: The assessment science team decides what we're measuring and how a skill turns into test content, they own construct definition (what "good at X" actually means), blueprinting (how an assessment is structured), and content authoring (writing the actual questions). You partner with them on how scoring methods serve that intent. The assessment tech team builds and maintains the scoring pipeline; you work embedded with them: writing specs, implementing and coordinating improvements, and validating that what ships meets the quality bar. You also work with data engineering on data availability. It's a small, cross-functional team, and you sit inside product decisions rather than in a separate research function.
WHAT YOU'LL OWN
Build trust by owning our scoring mechanism. These are the outcomes you’re accountable for:
- Own scoring quality end to end: evaluator design, rubric anchoring, calibration, and accountability for accuracy against expert benchmarks
- Build and run the continuous evaluation harness: gold sets, bias diagnostics, drift detection, and the pre-release gate that every scoring change must clear
- Define and publish assessment quality KPIs (human to AI agreement, reliability, classification accuracy, bias indicators, latency, cost per assessment) on dashboards anyone in the company can reference
- Lead the analytics investigations behind assessment decisions: performance studies, impact simulations,

← All remote jobs

Similar for you