Instead of applying uniform data processing rules, adapting the cleaning strategy per example—deciding what operation each piece of data needs—improves LLM pretraining efficiency and downstream performance.
DataOrchestra is a framework that customizes data processing for each example in pretraining, rather than applying one fixed strategy to all data. An orchestrator decides whether to drop, keep, or clean each data chunk, and if cleaning is needed, selects specific operations like editing or rewriting.