For the complete documentation index, see llms.txt. This page is also available as Markdown.

Dataset Curation

Learn how to get labeling recommendations using Tensorleap's Dataset Curation functionality.

After evaluation has completed, you can initiate 3 different curation process:

  1. Unlabeled Data Recommendations - Tensorleap analyzes your model's performance and data distribution to identify high-impact unlabeled samples. By focusing your labeling effort on the most informative samples, you can significantly accelerate model improvement and reduce labeling costs.

  2. Synthetic Data Optimization - Tensorleap aligns your synthetic data distribution with the target data source. By optimizing synthetic data generation, you can accelerate model improvement while significantly reducing labeling costs.

  3. Pruning - Tensorleap analyzes your dataset distribution and rebalances it using your selected metadata tags. Apply filters to focus on a subset and optionally prioritize specific metadata dimensions to guide the pruning process.

Last updated

Was this helpful?