Methods, systems, and media for providing direct and hybrid data acquisition approaches
Abstract:
Methods, systems, and media for providing direct and hybrid data acquisition approaches are provided. In accordance with some embodiments of the disclosed subject matter, a method of data acquisition for construction of classification models that incorporates multiple human reviewing resources is provided, the method comprising: receiving a cost structure for constructing a classification model using a data set; instructing a plurality of human reviewing resources to search through the data set and select one or more instances of a class that satisfy at least one criterion, wherein the plurality of human reviewing resources are provided with a definition of the class; training the classification model with the one or more instances from the plurality of human reviewing resources; determining when an expected gain for performing additional searches by the plurality of human reviewing resources as a function of the cost structure is lower than a given threshold; and, in response to determining that the expected gain as a function of the cost structure is lower than the given threshold, instructing the plurality of human reviewing resources that was searching through the data set to label one or more examples from the data set.
Information query
Patent Agency Ranking
0/0