Pearl Talent - Assessment Psychometrist

๐Ÿข Pearl Talent ยท all Pearl Talent jobs
๐Ÿ“ Mexico
๐Ÿ“… Posted 2026-08-16 ยท via Himalayas
๐Ÿท Psychometrics,Quantitative-Psychology,Assessment-Science,Psychological-Measurement,Research-Science,Psychology-Assessment-Expert,Psychometrician,Talent-Assessments,Psychometric-Assessment,Psychology-Assessment
Apply on original site โ†—
Pearl Talent - Psychometrician We're hiring a Psychometrician for Pearl, which finds exceptional talent from around the world, trains them to be AI-native, and places them into operational roles at startups as managed contractors: from client-facing roles to software engineers to executive assistants. We're 3x founders who've bootstrapped our company to a couple million in ARR and are adding six to seven figures in net new annualized revenue each month. Our clients span venture-backed tech and healthcare, including fast-growing startups and phenomenal US-based businesses that have raised over $3B in funding from Sequoia, a16z, Founders Fund, Y Combinator, and other top VC firms. Today we're roughly 50 people managing a few hundred talents, growing fast and into new verticals. We started Pearl because we believe that even though opportunity isn't created equal in the world, ambitious talent is. Location: Fully remote Purpose of Your Role You'll be the scientific backbone of our AI voice-model initiative. We're building a model that infers psychological constructs from voice, and its quality will only ever be as good as the measurement system behind it. This is not a machine-learning engineering role. You won't build the model itself. Instead, you'll define what we should measure, design and validate the instruments that measure it, and convert those measurements into defensible ground-truth labels and evaluation criteria for model training. You'll also be the person in the room with the authority โ€” and the responsibility โ€” to say what voice can't responsibly tell us about a person. What You'll Own 1. The Psychometric Framework - Review the literature and evaluate established frameworks (Big Five/OCEAN, HEXACO, DISC, and other relevant models) for validity, usefulness, and fit for our use case - Recommend which constructs and dimensions to include, exclude, or treat cautiously โ€” and explicitly identify traits that cannot be responsibly inferred from voice - Build and maintain a clear construct map connecting constructs, dimensions, indicators, survey items, and resulting scores, documented as the single source of truth for Research, Data, and AI teams 2. Instrument Design and Validation - Design psychometric surveys that serve as a high-quality measurement source for participants who also provide voice samples: items, response scales, scoring rules, reverse-coded items, and attention checks, while minimizing fatigue and response bias - Decide when to use validated existing scales, adapt them, or develop fit-for-purpose measures โ€” then pilot and iterate based on empirical performance - Run the full validation battery: reliability (McDonald's omega, Cronbach's alpha, item-total analysis), EFA and CFA to test dimensional structure, construct/convergent/discriminant/criterion validity, test-retest where appropriate, and IRT where useful - Recommend sample sizes, pilot methodology, and evidence thresholds a measure must clear before it's used for training 3. Ground Truth and Model Evaluation - Partner with AI and Data teams to transform psychometric responses into defensible training labels: continuous scores, categories, normalization, and confidence/reliability information - Define rules for missing, inconsistent, or low-quality responses, and the criteria for when a label is reliable enough to train on - Design the framework for comparing survey-based ground truth against voice-model predictions, and define evaluation metrics that reflect psychometric validity โ€” not just ML performance - Analyze where the model performs well, where it fails, and whether prediction quality differs by construct or population 4. Bias, Fairness, and Scientific Guardrails - Evaluate measurement invariance and potential bias across languages, cultures, and populations, including the impact of translation and sampling choices - Define the scientific limits around what conclusions may and may not be drawn from voice-based predi

โ† All remote jobs