Scene level video search
Abstract:
In some embodiments, a method trains a first prediction network to predict similarity between images in videos. The training uses boundaries detected in the videos to train the prediction network to predict images in a same scene to have similar feature descriptors. The first prediction network generates feature descriptors that describe library images from videos in a video library offered to users of a video delivery service. A search image is received and the prediction network predicts one or more library images for one or more videos that are predicted to be similar to the received image. The one or more library images for the one or more videos are provided as a search result.
Public/Granted literature
Information query
Patent Agency Ranking
0/0