Generating and applying data extraction templates
    11.
    发明授权
    Generating and applying data extraction templates 有权
    生成和应用数据提取模板

    公开(公告)号:US09563689B1

    公开(公告)日:2017-02-07

    申请号:US14470510

    申请日:2014-08-27

    Applicant: Google Inc.

    CPC classification number: G06F17/30705

    Abstract: Methods, apparatus, and computer-readable media are provided for generating and applying data extraction templates. In various implementations, a corpus of structured communications such as emails may be grouped into clusters based on one or more similarities between the structured communications. A set of structural paths may be identified from structured communications of a particular cluster. One or more structural paths of the set may be classified as transient wherein a count of occurrences of one or more associated segments of text across the particular cluster satisfies a criterion. One or more transient paths may be assigned a semantic data type and/or a confidentiality designation based on various signals. A data extraction template may be generated to extract, from subsequent structured communications, segments of text associated with transient (and in some cases, non-confidential) structural paths.

    Abstract translation: 提供了用于生成和应用数据提取模板的方法,装置和计算机可读介质。 在各种实现中,诸如电子邮件的结构化通信语料库可以基于结构化通信之间的一个或多个相似性被分组成群集。 可以从特定集群的结构化通信中识别一组结构路径。 该集合的一个或多个结构路径可以被分类为瞬时,其中跨越特定集群的一个或多个相关联的文本段的出现次数满足标准。 可以基于各种信号为一个或多个瞬态路径分配语义数据类型和/或机密性指定。 可以生成数据提取模板,以从后续结构化通信中提取与瞬态(以及在一些情况下,非机密)结构路径相关联的文本段。

Patent Agency Ranking