🐿️
4
c/ai-innovationsbrian303brian30318d agoProlific Poster

5 years ago I trained a model on my old blog posts, the output was weirdly familiar

Back in 2019 I scraped about 400 posts from my LiveJournal and ran them through a basic LSTM just for fun. The thing that stuck with me wasn't the gibberish, it was how it picked up my tics like ending sentences with 'or whatever' and overusing the word 'literally'. Now everyone talks about fine-tuning LLMs on personal data, but most people skip the cleaning step. My raw posts had typos, meme references, and inside jokes that made the output useless for anything real. What I learned is that garbage in means you get a chatbot that sounds like you after three beers, not like a thoughtful version of you. Has anyone else tried feeding their old social media or journal entries into a modern model and actually gotten something useful out of it?
0 comments

Log in to join the discussion

Log In
0 Comments

No comments yet

Be the first to share your thoughts on this discussion.