Abstract:In this thesis, feature of Chinese characters in regular fragments of document is studies and a method of extracting line-information of text is proposed. By defining the concept of L1-norm based differences between adjacent fragments, we develop a reassembly algorithm base on 0-1 programming and reduce the algorithm complexity by using cluster analysis. Compared with existing method, our reassembly method can fulfill the reassemble of given fragmented Chinese text effectively and efficiently without artificial supplementary.