Update: I stopped fine tuning my chatbot on 3,000 support tickets and started with 40
I used to believe more training data always meant a better bot, so I dumped 3,000 old support tickets into the model and spent two weeks cleaning them. The thing kept giving weird answers about refund policies from 2022 that we do not even honor now. Last month a guy named Dev at our Austin office told me to just pick the 40 cleanest tickets from the past 90 days and try that. The new version answered test questions correctly on the first try, and it took me one afternoon. So much of the advice online says scale up first, but my own logs say the opposite. Has anyone else had a small, fresh dataset beat a big messy one?