`Smith , IV et al .
`
`US 11,126,853 B2
`( 10 ) Patent No .:
`( 45 ) Date of Patent :
`Sep. 21 , 2021
`
`US011126853B2
`
`( 54 ) VIDEO TO DATA
`( 71 ) Applicant : CELLULAR SOUTH , INC . ,
`Ridgeland , MS ( US )
`( 72 ) Inventors : Bartlett Wade Smith , IV , Madison ,
`MS ( US ) ; Allison A. Talley , Ridgeland ,
`MS ( US ) ; John Carlos Shields , Fort
`Worth , TX ( US )
`( 73 ) Assignee : CELLULAR SOUTH , INC . ,
`Ridgeland , MS ( US )
`Subject to any disclaimer , the term of this
`patent is extended or adjusted under 35
`U.S.C. 154 ( b ) by 86 days .
`( 21 ) Appl . No .: 16 / 271,773
`( 22 ) Filed :
`Feb. 8 , 2019
`( 65 )
`
`( * ) Notice :
`
`( 56 )
`
`G06K 9/66 ( 2013.01 ) ; GIOL 15/26 ( 2013.01 ) ;
`H04N 21/23439 ( 2013.01 ) ; H04N 21/23608
`( 2013.01 ) ; H04N 21/8456 ( 2013.01 ) ; G06K
`2209/25 ( 2013.01 ) ; G06K 2209/27 ( 2013.01 )
`( 58 ) Field of Classification Search
`None
`See application file for complete search history .
`References Cited
`U.S. PATENT DOCUMENTS
`8/2006 Trivedi
`G06K 900241
`2006/0187305 A1 *
`348/169
`7/2010 Prokoski
`2010/0189313 A1 *
`A61B 5/411
`382/118
`2011/0305394 Al * 12/2011 Singer
`G06K 9/46
`382/190
`2/2015 Lakhani
`GIOL 25/57
`2015/0050010 A1 *
`386/285
`
`* cited by examiner
`Primary Examiner — Delomia L Gilliard
`( 74 ) Attorney , Agent , or Firm — Steptoe & Johnson LLP
`ABSTRACT
`( 57 )
`A method and system can generate video content from a
`video . The method and system can include a coordinator , an
`image detector , and an object recognizer . The coordinator
`can be communicatively coupled to a splitter and / or to a
`plurality of demultiplexer nodes . The splitter can be con
`figured to segment the video . The demultiplexer nodes can
`be configured to extract audio files from the video and / or to
`extract still frame images from the video . The image detec
`tor can be configured to detect images of objects in the still
`frame images . The object recognizer can be configured to
`compare an image of an object to a fractal . The recognizer
`can be further configured to update the fractal with the
`image . The coordinator can be configured to embed meta
`data about the object into the video .
`11 Claims , 17 Drawing Sheets
`
`Prior Publication Data
`Nov. 7 , 2019
`US 2019/0340437 A1
`Related U.S. Application Data
`Continuation of application No. 15 / 197,727 , filed on
`Jun . 29 , 2016 , now Pat . No. 10,204,274 .
`Int . CI .
`G06K 9/00
`GO6K 9/62
`G06F 40/40
`GIOL 15/26
`GO6K 9/66
`H04N 21/2343
`H04N 21/236
`HO4N 21/845
`U.S. CI .
`CPC
`
`( 2006.01 )
`( 2006.01 )
`( 2020.01 )
`( 2006.01 )
`( 2006.01 )
`( 2011.01 )
`( 2011.01 )
`( 2011.01 )
`
`GOOK 9/00718 ( 2013.01 ) ; G06F 40/40
`( 2020.01 ) ; G06K 9/00201 ( 2013.01 ) ; G06K
`9/00261 ( 2013.01 ) ; G06K 9/6256 ( 2013.01 ) ;
`
`( 63 )
`
`( 51 )
`
`( 52 )
`
`Wirles
`??? ?? ???
`
`
`
`
`Cato
`
`Muotik
`
`SAYAW
`
`Litty .
`Pro
`
`dywuls
`F.COM
`
`UTAMA
`re
`
`CarWw
`
`???????
`TTCL
`True !
`
`megin
`
`|
`
`??
`
`??
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 001
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 1 of 17
`
`US 11,126,853 B2
`
`Video
`
`110
`
`Audio to text
`
`Image to text
`
`140
`
`150
`
`1
`
`Natural
`language
`processing
`
`1
`Natural
`language
`processing
`
`120
`
`130
`
`Combine image text
`and audio text
`
`160
`
`Generate
`video text
`
`Figure 1
`
`170
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 002
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 2 of 17
`
`US 11,126,853 B2
`
`Server or servers 220
`
`Network 230
`
`User equipment 210
`
`Figure 2
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 003
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 3 of 17
`
`US 11,126,853 B2
`
`Video data
`310
`
`Distributed image
`data processing
`320-1
`
`Distributed image
`data processing
`320-2
`
`Distributed image
`data processing
`320 - N
`
`II /
`
`Combine 330
`
`Figure 3
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 004
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 4 of 17
`
`US 11,126,853 B2
`
`Audio data
`410
`
`Distributed audio
`data processing
`420-1
`
`Distributed audio
`data processing
`420-2
`
`Distributed audio
`data processing
`420 - N
`
`al / /
`
`Combine 430
`
`Figure 4
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 005
`
`
`
`U.S. Patent
`
`Sep.21 , 2021
`
`Sheet 5 of 17
`
`US 11,126,853 B2
`
`Output M
`
`-C
`
`> ?
`OM
`* C
`* C
`? ? ? COS ?
`
`| ---
`
`Figure 5
`
`? e
`wwwwwwwwwwww
`
`? ??
`
`?
`
`Ange
`- |
`
`Marie
`
`Sys JS
`
`32 ? 3V
`
`.
`
`}
`Image
`
`Audio File
`
`Data
`
`? ?
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 006
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 6 of 17
`
`US 11,126,853 B2
`
`Figure 6
`
`385.VOLIX
`
`19K / Quinot
`
`Proxies
`
`Distrileke Aa
`
`Distributes Ave
`
`Sasagne
`
`Furder MLA
`
`Sainte Topics
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 007
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 7 of 17
`
`US 11,126,853 B2
`
`06L
`
`S61
`
`Figure 7
`
`730
`
`Ob
`
`merce
`
`720
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 008
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 8 of 17
`
`US 11,126,853 B2
`
`Boundary
`
`
`
`Mark with
`
`
`
`Filtering Uso
`
`
`
`Eigen Vector
`
`Erameters
`
`woon
`
`
`
`
`
`Set Create Elgen
`
`Dette
`
`Figure 8
`
`
`
`
`
`Retetexte Datis Sent
`
`BASES
`
`image ? mes
`
`Begins
`
`008
`
`? ?? ? ?
`
`810
`
`
`
`Image Detection
`
`
`
`Image Recognition
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 009
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 9 of 17
`
`US 11,126,853 B2
`
`Figure 9
`
`us
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0010
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 10 of 17
`
`US 11,126,853 B2
`
`Aggregate
`Recognition
`
`Fractal Located
`
`Items
`
`Training
`
`Set
`
`Figure 10
`
`Segments
`
`Split
`
`Coordinator
`Demux
`
`Media
`
`Object
`
`Face
`
`Text
`
`Logo
`
`Dialog
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0011
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 11 of 17
`
`US 11,126,853 B2
`
`End Pre - processing
`
`still Images
`into Spits Meda Segment
`
`
`
`
`
`Demultiplexer Node
`
`Processing is Complete
`
`If Available , Coordinator sends Additional
`
`image
`
`
`
`Demultiplexor Mode
`
`Distributa
`
`Media = Segment to Demultiplexer
`
`Processing
`
`Figure 11
`
`Optimal Processing Configurations
`Asset Alcibutes to Datamine
`Analyzes
`
`
`
`Splitter Component
`
`De multiplexer
`
`Recognizes Processing
`Request
`
`Uploaded to System
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0012
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 12 of 17
`
`US 11,126,853 B2
`
`image Frepressing
`Carpet
`
`ises Friesina
`Tastami
`
`Trairing
`
`centry Fracercing
`
`Ener
`
`For Teri
`anima
`
`scerary Training
`
`?? . ????
`Training Prograine
`
`attributes
`
`-
`
`Cataractice :
`for Image
`
`For Ever
`An image
`
`?????
`
`Compare
`
`Refere
`Latest
`
`? ra recognition :
`Faceritage
`
`it " -3 - ata in
`First
`
`???????
`match Extra
`Baits for
`
`Store !
`Fire ants
`
`For ordinancer
`Toorate
`
`???? itical
`ma
`Frames
`
`Figure 13
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0013
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 13 of 17
`
`US 11,126,853 B2
`
`Processing Complete
`
`System lo entifies
`Possible a Matches
`
`High Confidence
`
`Image Data Rzeged
`For Analysis
`
`Image Data added to
`
`30 Rotated
`Image
`
`************
`
`Original Media aset
`RePredested
`
`End
`
`Figure 14
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0014
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 14 of 17
`
`US 11,126,853 B2
`
`Processing Complete
`
`Aarget Media
`Type Supports
`
`Create Sustitie SRT
`Compatible Copy of
`
`Trasfomed into
`
`Date Stream
`
`Original Media 49et :
`
`Embed Meta - Data
`
`Figure 15
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0015
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 15 of 17
`
`US 11,126,853 B2
`
`C 100 % 2383
`
`Q Sexkoos
`OOO
`
`???????? 1 ?????
`
`Figure 16
`
`Frames : 169 of 201
`
`P2
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0016
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 16 of 17
`
`US 11,126,853 B2
`
`W ? ES
`
`
`
`w ASUS ****
`
`KRESS
`
`Figure 17
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0017
`
`
`
`U.S. Patent
`
`Sep. 21 , 2021
`
`Sheet 17 of 17
`
`US 11,126,853 B2
`
`YO
`
`KRESS : ?? MONY
`
`Wood
`
`3
`
`Figure 18
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0018
`
`
`
`5
`
`10
`
`15
`
`US 11,126,853 B2
`
`1
`VIDEO TO DATA
`
`CLAIM OF PRIORITY
`This application is a continuation of U.S. application Ser .
`No. 15 / 197,727 , filed Jun . 29 , 2016 , now U.S. Pat . No.
`10,204,274 , which is incorporated by reference in its
`entirety .
`
`2
`into video segments , extracting an audio file from a segment
`of the video segments , extracting a video frame file of still
`frames from the segment , detecting an image of an object in
`the still frames , recognizing the object as a specific object ,
`updating an object - specific fractal with the image , and
`embedding metadata in the video about the specific object .
`In some embodiments , the metadata can include a time
`stamp and / or a coordinate location of the object in one or
`more of the still frames . The metadata can include a recog
`TECHNICAL FIELD
`nition confidence score . The method can further include
`distributing the video segments across a plurality of proces
`The present invention relates to a method and a system for
`sors . The method can include extracting a plurality of video
`generating various and useful data from source media , such
`frame files , such as all of the video segments , by a plurality
`as videos and other digital content . The data can be embed
`of parallel processors .
`ded within the source media or combined with the source
`In other embodiments , the video can be a stereoscopic
`media for creating an augmented video containing additional
`three - dimensional video .
`contextual information .
`In yet other embodiments , the method can include gen
`erating text based on extracted audio file and / or applying
`BACKGROUND
`20 natural language processing to the text . The method can
`include determining context associated with the video based
`In the field of image contextualization , distributed reverse
`on the natural language processing .
`image similarity searching can be used to identify images
`In some embodiments , the method can include processing
`similar to a target image . Reverse image searching can find
`the video frame file to extract image text . The object can be
`exactly matching images as well as flipped , cropped , and
`altered versions of the target image . Distributed reverse 25 a face or a logo . The object can be recognized as a three
`image similarity searching can be used to identify symbolic
`dimensional rotation of a known object .
`similarity within images . Audio - to - text algorithms can be
`In other embodiments , a three - dimensional fractal can be
`used to transcribe text from audio . An exemplary application
`updated , e.g. , with the image of the object . The method can
`is note - taking software . Audio - to - text , however , lacks
`include generating a content - rich video based on the video
`semantic and contextual language understanding .
`30 and the metadata .
`Another aspect can include a system for generating data
`from a video . The system can include a coordinator , an
`SUMMARY
`image dete
`and an object recognizer . The coordinator
`The present invention is generally directed to a method to
`can be communicatively coupled to a splitter and / or to a
`generate data from video content , such as text and / or image- 35 plurality of demultiplexer nodes . The splitter can be con
`related information . A server executing the method can be
`figured to segment the video . The demultiplexer nodes can
`directed by a program stored on a non - transitory computer-
`be configured to extract audio files from the video and / or to
`readable medium . The video text can be , for example , a
`extract still frame images from the video . The image detec
`tor can be configured to detect images of objects in the still
`context description of the video .
`An aspect can include a system for generating data from 40 frame images . The object recognizer can be configured to
`a video . The system can include a coordinator , an image
`compare an object image of an object to a fractal . The
`detector , and an object recognizer . The coordinator can be
`recognizer can be further configured to update the fractal
`communicatively coupled to a splitter and / or to a plurality of
`with the object image . The coordinator can be configured to
`demultiplexer nodes . The splitter can be configured to
`generate one or more metadata streams corresponding to the
`segment the video . The demultiplexer nodes can be config- 45 images . The one or more metadata streams can include
`ured to extract audio files from the video and / or to extract
`timestamps corresponding to the images . The coordinator
`still frame images from the video . The image detector can be
`can be configured to embed the metadata streams in the
`configured to detect images of objects in the still frame
`video .
`images . The object recognizer can be configured to compare
`In some embodiments , the metadata streams can be
`an image of an object to a fractal . The recognizer can be 50 embedded in the video as subtitle resource tracks .
`further configured to update the fractal with the image . The
`In other embodiments , the system can be accessible over
`coordinator can be configured to embed metadata about the
`a network via application program interfaces ( APIs ) .
`object into the video .
`In yet other embodiments , the coordinator can be further
`In some embodiments , the metadata can include a time-
`configured to output the video according to multiple video
`stamp and / or a coordinate location of the object in one or 55 formats . For example , the coordinator can be configured to
`more of the still frame images . The coordinator can be
`automatically generate data files in a variety of formats for
`configured to create additional demultiplexer processing
`delivery independent of the video . The system in some
`capacity . The coordinator can be configured to create addi-
`embodiments can embed data as a stream , as a wrapper ,
`tional demultiplexer nodes , e.g. , when the demultiplexer
`and / or as a subtitle resource track . The coordinator can be
`nodes reach at least 80 % of processing capacity .
`60 configured to read / write to / from Media Asset Management
`In other embodiments , the demultiplexer nodes can gen-
`Systems , Digital Asset Management Systems , and / or Con
`erate a confidence score based on a comparison of the image
`tent Management Systems .
`and the fractal . In yet other embodiments , the recognizer can
`In some embodiments , the system can be configured to
`generate a confidence score based on a comparison of the
`capture the geolocation of objects in a video . The system can
`65 be configured to derive a confidence score for each instance
`image and the fractal .
`Another aspect can include a method to generate data
`of recognition . The system can be configured to apply
`from a video . The method can include segmenting the video
`natural language processing , for example , for associative
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0019
`
`
`
`US 11,126,853 B2
`
`3
`4
`FIG . 18 depicts a distorted image after calibration accord
`terms and / or to apply contextual analysis of corresponding
`ing to present embodiments .
`data points ( such as audio , objects , etc. ) to verify accuracy .
`An aspect can include a method of creating data from a
`DETAILED DESCRIPTION
`video by machine recognition . The method can include
`extracting an audio file from the video , segmenting the video 5
`A detailed explanation of the system and method accord
`into video frames of still images , distributing the video
`ing to exemplary embodiments of the present invention are
`segments to N processors , wherein N is an integer greater
`described below . Exemplary embodiments described ,
`than one , generating a timestamped transcript from the audio
`shown , and / or disclosed herein are not intended to limit the
`file , associating the timestamped transcript with correspond-
`ing video frames , deriving topics from the audio file based 10 claims , but rather , are intended to instruct one of ordinary
`on natural language processing , recognizing an object from
`skill in the art as to various aspects of the invention . Other
`still images , using a reference database to identify the object ,
`embodiments can be practiced and / or implemented without
`and embedding , within the video , data based on a recognized
`departing from the scope and spirit of the claimed invention .
`object , the topics , and the timestamped transcript .
`The present invention is generally directed to system ,
`In some embodiments , the video can be a virtual reality 15 device , and method of generating data from source media ,
`video file or a traditional video vile . Data based on the
`such as images , video , and audio . Video can include two
`recognized object can include a geolocation .
`dimensional video and / or stereoscopic three - dimensional
`In other embodiments , the method can include generating
`video such as virtual reality ( VR ) files . The generated data
`a plurality of video files . Each of the video files can include
`the video and the embedded data . Each of the plurality of 20 can include text and information relating to context , sym
`bols , brands , features , objects , faces and / or topics found in
`video files can be generated in a different format .
`the source media . In an embodiment , the video - to - data
`In other embodiments , the method can include generating
`a confidence score . The score can be associated with the
`engine can perform the functions directed by programs
`recognized object . The method can include analyzing the
`stored in a computer - readable medium . That is , the embodi
`25 ments can include hardware ( such as circuits , processors ,
`still images to determine context of the video .
`memory , user and / or hardware interfaces , etc. ) and / or soft
`ware ( such as computer - program products that include com
`DESCRIPTION OF THE DRAWINGS
`puter - useable instructions embodied on one or more com
`puter - readable media ) .
`The present invention is further described in the detailed
`The various video - to - data techniques , methods , and sys
`description which follows , in reference to the noted plurality 30
`of drawings by way of non - limiting examples of certain
`tems described herein can be implemented in part or in
`embodiments of the present invention , in which like numer-
`whole using computer - based systems and methods . Addi
`als represent like elements throughout the several views of
`tionally , computer - based systems and methods can be used
`the drawings , and wherein :
`to augment or enhance the functionality described herein ,
`FIG . 1 illustrates an exemplary workflow in certain 35 increase the speed at which the functions can be performed ,
`and provide additional features and aspects as a part of , or
`embodiments .
`FIG . 2 illustrates an embodiment of image data process-
`in addition to , those described elsewhere herein .
`ing .
`Various computer - based systems , methods , and imple
`FIG . 3 illustrates aspects of image data processing .
`mentations in accordance with the described technology are
`40 presented below .
`FIG . 4 illustrates aspects of audio data processing .
`FIG . 5 illustrates various exemplary aspects of embodi-
`A video - to - data engine can be embodied by a computer or
`ments of the present invention .
`a server and can have an internal or external memory for
`FIG . 6 illustrates a flow diagram of a present embodiment
`storing data and programs such as an operating system ( e.g. ,
`FIG . 7 illustrates exemplary architecture of a present
`DOS , Windows2000TM , Windows XPTM , Windows NTTM ,
`45 OS / 2 , UNIX , Linux , Xbox OS , Orbis OS , and FreeBSD )
`embodiment .
`FIG . 8 illustrates a flow diagram of an embodiment of
`and / or one or more application programs . The video - to - data
`image recognition .
`engine can be implemented by a computer or a server
`FIG . 9 illustrates an embodiment of a graphical user
`through tools of a particular software development kit
`( SDK ) . Examples of application programs include computer
`interface of the present invention .
`FIG . 10 illustrates exemplary system architecture with an 50 programs implementing the techniques described herein for
`exemplary process flow .
`lyric and multimedia customization , authoring applications
`FIG . 11 illustrates an exemplary process for distributed
`( e.g. , word processing programs , database programs , spread
`demultiplexing and preparation of source media files .
`sheet programs , or graphics programs ) capable of generating
`FIG . 12 illustrates exemplary distributed processing and
`documents , files , or other electronic content ; client applica
`aggregation .
`55 tions ( e.g. , an Internet Service Provider ( ISP ) client , an
`FIG . 13 illustrates an exemplary process for improved
`e - mail client , or an instant messaging ( IM ) client ) capable of
`communicating with other computer users , accessing vari
`recognition based on near frame proximity .
`FIG . 14 illustrates an exemplary process for improved
`ous computer resources , and viewing , creating , or otherwise
`recognition based on partial three - dimensional matching .
`manipulating electronic content ; and browser applications
`FIG . 15 illustrates an exemplary process for embedding 60 ( e.g. , Microsoft's Internet Explorer ) capable of rendering
`extracted data to original source files as metadata .
`standard Internet content and other content formatted
`FIG . 16 depicts an exemplary interface showing a 360 °
`according to standard protocols such as the Hypertext Trans
`image from a virtual reality video file and embedded meta-
`fer Protocol ( HTTP ) . One or more of the application pro
`grams can be installed on the internal or external storage of
`data .
`FIG . 17 is an image of the Kress Building in Ft . Worth 65 the computer . Application programs can be externally stored
`in or performed by one or more device ( s ) external to the
`Tex . as taken by a fisheye lens , as used in virtual reality
`images .
`computer .
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0020
`
`
`
`US 11,126,853 B2
`
`5
`6
`separate tracks or as a single track . Text extraction can be
`The computer or server can include a central processing
`optimized by utilizing audio segments of various lengths in
`unit ( CPU ) for executing instructions in response to com-
`time . For example , if a segment of audio is greater than one
`mands , and a communication device for sending and receiv-
`minute , the engine can split the audio track in half . In this
`ing data . One example of the communication device can be
`a modem . Other examples include a transceiver , a commu- 5 case , the engine can first analyze that specific sequence for
`nication card , a satellite dish , an antenna , a network adapter ,
`dialog at the timestamp of the potential split . If the segment
`or some other mechanism capable of transmitting and
`at the split contains audio , the system can split the audio at
`receiving data over a communications link through a wired
`the next silent point in the track to avoid splitting tracks
`or wireless data pathway .
`mid - word . Each segment is processed using the Kaldi pro
`The computer or server can also include an input / output 10 cess for speech recognition and dialog extraction . Segments
`interface that enables wired or wireless connection to vari-
`can be subsequently processed through , for example , LIUM
`ous peripheral devices . In one implementation , a processor-
`speaker diarization . Results can be applied to a result
`based system of the computer can include a main memory ,
`datastore for analysis or later processing .
`preferably random access memory ( RAM ) , and can also
`An example of the image data processing is illustrated in
`include
`secondary memory , which can be a tangible 15 FIG . 3. The video - to - data engine can segment the video into
`computer - readable medium . The tangible computer - read-
`chunks for distributed , or parallel , processing as shown
`able medium memory can include , for example , a hard disk
`schematically in FIG . 3. Distributed processing in this
`drive or a removable storage drive , a flash based storage
`context can mean that the processing time for analyzing a
`system or solid - state drive , a floppy disk drive , a magnetic
`video from beginning to end is a fraction of the play time of
`tape drive , an optical disk drive ( Blu - Ray , DVD , CD drive ) , 20 the video . This can be accomplished by breaking the pro
`magnetic tape , paper tape , punched cards , standalone RAM cesses into sections and processing them simultaneously .
`disks , Iomega Zip drive , etc. The removable storage drive
`The images and audio can each be broken up into pieces
`can read from or write to a removable storage medium . A such that the meaning of a continuous message is preserved .
`removable storage medium can include a floppy disk , mag-
`At 120 , the video - to - data engine performs an image data
`netic tape , optical disk ( Blu - Ray disc , DVD , CD ) a memory 25 processing on the video stream . In FIG . 3 , the image data
`card ( CompactFlash card , Secure Digital card , Memory
`310 can be segmented into N segments and processed in
`Stick ) , paper data storage ( punched card , punched tape ) , etc. ,
`parallel ( e.g. , distributed processing 320-1 to 320 - N ) , allow
`which can be removed from the storage drive used to
`ing for near real - time processing .
`perform read and write operations . As will be appreciated ,
`An example of the video image data processing can be
`the removable storage medium can include computer soft- 30 symbol ( or object ) based . Using an image processing tech
`nique such as color edge detection , a symbol of a screen or
`ware or data .
`In alternative embodiments , the tangible computer - read-
`an image of the video can be isolated . The symbol can be
`able medium memory can include other similar means for
`identified using an object template database . For example ,
`allowing computer programs or other instructions to be
`the symbol includes 4 legs and a tail , and when matched with
`loaded into a computer system . Such means can include , for 35 the object template database , the symbol may be identified
`example , a removable storage unit and an interface .
`as a dog . The object template database can be adaptive and
`Examples of such can include a program cartridge and
`therefore , the performance would improve with usage .
`cartridge interface ( such as found in video game devices ) , a
`Other image data processing techniques can include
`removable memory chip ( such as an EPROM or flash
`image extraction , high - level vision and symbol detection ,
`memory ) and associated socket , and other removable stor- 40 figure - ground separation , depth and motion perception .
`age units and interfaces , which allow software and data to be
`These and / or other image data processing techniques can be
`transferred from the removable storage unit to the computer
`utilized to build a catalogue and / or repository of extracted
`objects . Recognized information about the extracted
`system .
`An embodiment of video - to - data engine operation is
`objects such as object type , context , brands / logos , vehicle
`illustrated in FIG . 1. At 110 , a video stream is presented . The 45 type / make / model , clothing worn , celebrity name , etc.
`can
`video stream can be in one or more of the formats ( but not
`be added to a file associated with the extracted object and / or
`limited to ) : Advanced Video Codec High Definition
`used to augment metadata in the video from which the image
`( AVCHD ) , Audio Video Interlaced ( AVI ) , Flash Video For-
`data was processed .
`mat ( FLV ) , Motion Picture Experts Group ( MPEG ) , Win-
`Another example of video image processing can be color
`dows Media Video ( WMV ) , or Apple QuickTime ( MOV ) , 50 segmentation . The colors of an image ( e.g. , a screen ) of the
`video can be segmented or grouped . The result can be
`h.264 ( MP4 ) .
`The engine can extract audio data and image data ( e.g.
`compared to a database using color similarity matching .
`images or frames forming the video ) from the video stream .
`Based on the identified symbol , a plurality of instances of
`The engine can detect and identify objects , faces , logos , text ,
`the symbol can be compared to a topic database to identify
`music , sounds and spoken language in video by means of 55 a topic ( such as an event ) . For example , the result may
`demultiplexing and extracting features from the video and
`identify the dog ( symbol ) as running or jumping . The topic
`passing those features into a distributed system as the video
`database can be adaptive to improve its performance with
`loads into the network I / O buffer . In some embodiments , the
`usage .
`video stream and the extracted image data can be stored in
`Thus , using the processing example above , text describing
`a memory or storage device such as those discussed above . 60 a symbol of the video and topic relating to the symbol can
`A copy of the extracted image data can be used for process-
`be generated , as is illustrated in FIG . 9. Data generated from
`ing .
`an image and / or from audio transcription can be time
`The system and method can include dialog extraction .
`stamped , for example , according to when it appeared , was
`Language and vocabulary models can be included in the
`heard , and / or according to the video frame from which it
`system to support desired languages . Multiple languages can 65 was pulled . The time - stamped data can be physically asso
`be incorporated into the system and method . The engine can
`ciated with the video as metadata embedded at the relevant
`process audio media containing multiple audio tracks as
`portion of video .
`
`Google Exhibit 1001 - Google v. CSI
`IPR2025-00877 - Page 0021
`
`
`
`US 11,126,853 B2
`
`7
`8
`nition of the player's jersey number , and by cross referenc
`At 330 , the engine combines the topics as an array of keys
`ing this data with that team's roster ( as opposed to another
`and values with respect to the segments . The engine can
`team , which is an example of why the logo recognition can
`segment the topics over a period of time and weight the
`be important ) . Such embodiments can further learn to iden
`strength of each topic . Further , the engine applies the topical
`metadata to the original full video . The image topics can be 5 tify that player more readily and save his image as data .
`stored as topics for the entire video or each image segment .
`Similarly , the audio transcript of a video can be used to
`The topic generation process can be repeated for all identi-
`derive certain context helpful in identifying and correcting
`fiable symbols in a video in a distributed process . The
`or eliminating image false positives . In this way , an image
`outcome would be several topical descriptors of the content
`anomaly or anomalies identified in a given video frame ( s )
`within a video . An example of the aggregate information that 10 are associated with time ( time stamped ) and correlated with
`can be derived using the above example would be a deter-
`a time range from the transcribed audio to establish certain
`mination that the video presented a dog , which was jumping ,
`probabilities of accuracy .
`Moreover , the aforementioned methodologies establish
`on the beach , with people , by a resort .
`Although further described herein , image detection can be
`ing probabilities of accuracy of image identification from a
`considered a process of determining if a pattern or patterns 15 set of frames and from the audio transcription can be
`exist in an image and whether the pattern or patterns meet
`combined to improve the results . Improved results can be
`criteria of a face , image , and / or text . If the result is positive ,
`embedded in a video or an audio file and / or an image
`image recognition can be employed . Image recognition can
`file as metadata , as can probabilities of accuracy .
`generally be considered matching of detected objection to
`In some embodiments , a similar context methodology can
`known objects and / or matching through machine learning . 20 be used to identify unknown objects in a given image by
`Generally speaking , detection and recognition , while shar-
`narrowing a large , or practically infinite , number of possi
`bilities to a relatively small number of object po



