Natural Language Measurement Specialist
Natural Language Specialist, Educational Measurement & AI
College Board – Learning & Assessment
Location:
-
This is a remote role. Candidates who live near CB offices have theoptionof being fully remote or hybrid (Tuesday and Wednesday in office). All CB employeesare required tooccasionally travel to meet in person for business purposes.
Role Type :
- This is a full-time position
About the Team
The Automated Scoring team provides critical insights and tools to support the design, delivery, and continuous improvement of digital assessments. We operate at the intersection of educational measurement, data science, and emerging AI technologies. Our work spans the measurement of language-based constructs, the development of large language model (LLM) systems for feedback generation and annotation, and research to ensure the validity, fairness, and reliability of our systems.
We are a collaborative, mission-driven team that values both psychometric rigor and technical expertise. We combine modern machine learning approaches, including large language models, with strong measurement principles to create scalable, trustworthy solutions that expand opportunity for students.
About the Opportunity
As a Natural Language Specialist, you will help define and advance how language-based performance is measured in high-stakes educational settings. This role sits at the intersection of natural language processing, large language models, and educational measurement, and is ideal for someone who pairs strong measurement training with a working knowledge of modern AI.
You will translate measurement constructs into LLM-based feedback and annotation systems, design and conduct the studies that establish their validity, fairness, and reliability, and ensure that what our models produce holds up to rigorous psychometric standards.You will serve as the measurement authority for cross-functional partners—psychometricians, engineers, and data scientists—shaping how language-based constructs are defined, evaluated, and applied across theteam's portfolio of systems.Your primary lens will be measurement: defining what feedback and annotations mean, evidencing that they mean it, and improving them over time.
In this role, you will:
Natural Language Measurement & Psychometrics (40%)
-
Define and operationalize language-based constructs for automated annotation and feedback generation
-
Apply psychometric principles—reliability, validity, dimensionality, and measurement invariance—to LLM-based feedback and annotation systems
-
Design and lead validity studies, including human–machine agreement, rater comparison, and fairness analyses across subgroups
-
Develop and apply methods for detecting and mitigating bias in language-based scores
-
Establish/Recommend annotation guidelines, feedback quality criteria, and standards for acceptable model performance
-
Translate measurement requirements into specifications that guide model development and evaluation
LLM & AI Development (30%)
-
Contribute to prompt design, fine-tuning, and evaluation of LLM-based feedback and annotation systems
-
Develop and refine machine learning models for measuring language-based constructs
-
Build evaluation frameworks that connect model behavior to measurement outcomes
-
Collaborate with senior team members to translate measurement findings into production systems
-
Implement high-quality, maintainable code for model development and evaluation
Research & Validation (10%)
-
Lead and contribute to research studies that evaluate model performance and support assessment validity
-
Apply statistical and psychometric methods to analyze results and inform model improvements
-
Document methodologies and findings in a clear and rigorous manner
-
Stay current with advances in educational measurement, NLP, and learning science
Data Engineering & Pipelines (10%)
-
Prepare and curate datasets f