In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get. What you’ll do Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds Explore open source datasets and create internal ones most suitable to build fundamental World Models... Read the full ad