Invention Grant
US07805289B2 Aligning hierarchal and sequential document trees to identify parallel data 失效
对齐层次和顺序文档树以识别并行数据

Aligning hierarchal and sequential document trees to identify parallel data
Abstract:
A set of candidate parallel pages is identified based on trigger words in one or more pages downloaded from a given network location (such as a website). A set of document trees representing each of the candidate pages are aligned to identify translationally parallel content and hyperlinks. The parallel content is further fed into conventional sentence aligner for parallel sentences. And the parallel hyperlinks usually refer to other parallel documents, and lead to a recursive mining of parallel documents.
Information query
Patent Agency Ranking
0/0