Ascribing a confidence factor for identifying a given column in a structured dataset belonging to a particular sensitive type
Abstract:
Various aspects of the subject technology relate to systems, methods, and machine-readable media for determining a confidence factor for a sensitive type. The method includes applying a set of matching procedures to cells in a structured data set, the structured data set comprising columns and/or rows. The method also includes counting hit counts for the cells, the hit counts corresponding to successful matches. The method also includes counting null counts for the cells, the null counts corresponding to cells having null or invalid values. The method also includes counting mishit counts for the cells, the mishit counts corresponding to cells that are not null and do not result in a match. The method also includes calculating the confidence factor based on the hit counts, the null counts, and the mishit counts, the confidence factor providing an effective probability that cells in the structured data set is of the sensitive type.
Information query
Patent Agency Ranking
0/0