Invention Grant
- Patent Title: Implicit relation induction via purposeful overfitting of a word embedding model on a subset of a document corpus
-
Application No.: US15928310Application Date: 2018-03-22
-
Publication No.: US10885082B2Publication Date: 2021-01-05
- Inventor: Anastas Stoyanovsky , Roxana Gheorghiu , Robert L. Yates
- Applicant: International Business Machines Corporation
- Applicant Address: US NY Armonk
- Assignee: International Business Machines Corporation
- Current Assignee: International Business Machines Corporation
- Current Assignee Address: US NY Armonk
- Agency: Law Office of Jim Boice
- Main IPC: G06N20/00
- IPC: G06N20/00 ; G06F16/33 ; G06F16/93 ; G06F16/951

Abstract:
A method overfits a word vector generating process to identify implicit relationships between two or more terms in a corpus. A server identifies instances of multiple user-generated pairs of terms in an original corpus of documents, in which the terms are labeled but a relationship between two or more of the corpus terms are not identified. The server then extracts sentences, from the original corpus of documents, that contain one or more of the multiple user-generated pairs of terms, and combines the sentences into a training corpus, which is used to purposely overfit a word embedding model. This word embedding model leads to a vector that is used to identify other terms that have a same type of relationship as that found in the multiple user-generated pairs of terms, such that search corpus of documents can be searched for similar terms that trained the word embedding model.
Public/Granted literature
Information query