r/MachineLearning
· Communities
My OCR model mislabels section titles as body text. Is a CRF the right fix, or am I overcomplicating it? [P]
Hi everyone, I'm working on extracting the hierarchical structure of long PDF documents (legal/regulatory text, lots of numbered sections) and would like to gather some feedback on my approach before committing to it. What I've done so far: I render each PDF page to an image and run it through Baidu's DeepSeek-OCR mode