Invention Grant
- Patent Title: Preprocessing of string inputs in natural language processing
-
Application No.: US15376923Application Date: 2016-12-13
-
Publication No.: US10372816B2Publication Date: 2019-08-06
- Inventor: Charles E. Beller , Chengmin Ding , Allen Ginsberg , Elinna Shek
- Applicant: International Business Machines Corporation
- Applicant Address: US NY Armonk
- Assignee: International Business Machines Corporation
- Current Assignee: International Business Machines Corporation
- Current Assignee Address: US NY Armonk
- Agency: Lieberman & Brandsdorfer, LLC
- Main IPC: G06F17/27
- IPC: G06F17/27

Abstract:
Natural language processing of raw text data for optimal sentence boundary placement. Raw text is extracted from a document and subject to cleaning. The extracted raw text is examined to identify preliminary sentence boundaries, which are used to identify potential sentences in the raw text. One or more potential sentences are assigned a well-formedness score. A value of the score correlates to whether the potential sentence is a truncated/ill-formed sentence or a well-formed sentence. One or more preliminary sentence boundaries are optimized depending on the value of the score of the potential sentence(s). Accordingly, the processing herein is an optimization that creates a sentence boundary optimized output.
Public/Granted literature
- US20180165270A1 Preprocessing of String Inputs in Natural Language Processing Public/Granted day:2018-06-14
Information query