`
`Smith, IV et al.
`
`US011126853B2
`
`US 11,126,853 B2
`Sep. 21, 2021
`
`(10) Patent No.:
`45) Date of Patent:
`
`(54)
`
`(71)
`
`(72)
`
`(73)
`
`)
`
`@1
`
`(22)
`
`(65)
`
`(63)
`
`(1)
`
`(52)
`
`VIDEO TO DATA
`
`Applicant: CELLULAR SOUTH, INC.,
`Ridgeland, MS (US)
`
`Bartlett Wade Smith, IV, Madison,
`MS (US); Allison A. Talley, Ridgeland,
`MS (US); John Carlos Shields, Fort
`Worth, TX (US)
`
`Inventors:
`
`Assignee: CELLULAR SOUTH, INC.,
`
`Ridgeland, MS (US)
`
`Notice: Subject to any disclaimer, the term of this
`
`patent is extended or adjusted under 35
`U.S.C. 154(b) by 86 days.
`
`Appl. No.: 16/271,773
`Filed: Feb. 8, 2019
`
`Prior Publication Data
`
`US 2019/0340437 Al Nov. 7, 2019
`
`Related U.S. Application Data
`
`Continuation of application No. 15/197,727, filed on
`Jun. 29, 2016, now Pat. No. 10,204,274.
`
`Int. CL.
`
`GO6K 9/00 (2006.01)
`GO6K 9/62 (2006.01)
`GO6F 40/40 (2020.01)
`GI0L 15726 (2006.01)
`GO6K 9/66 (2006.01)
`HO4N 21/2343 (2011.01)
`HO4N 21236 (2011.01)
`HO4N 21/845 (2011.01)
`U.S. Cl
`
`CPC ........ GO6K 9/00718 (2013.01); GOGF 40/40
`
`(2020.01); GO6K 9/00201 (2013.01); GO6K
`9/00261 (2013.01); GO6K 9/6256 (2013.01);
`
`GO6K 9/66 (2013.01); G10L 15/26 (2013.01);
`HO4N 21/23439 (2013.01); HO4N 21/23608
`(2013.01); HO4N 21/8456 (2013.01); GO6K
`2209/25 (2013.01); GO6K 2209/27 (2013.01)
`(58) Field of Classification Search
`None
`See application file for complete search history.
`
`(56) References Cited
`
`U.S. PATENT DOCUMENTS
`
`2006/0187305 Al* 82006 Trivedi ............. GO6K 9/00241
`348/169
`2010/0189313 Al* 7/2010 Prokoski ................ A61B 5/411
`382/118
`2011/0305394 A1* 12/2011 Singer .........cc.... GO6K 9/46
`382/190
`2015/0050010 A1* 2/2015 Lakhani ... G10L 25/57
`386/285
`
`* cited by examiner
`
`Primary Examiner — Delomia L Gilliard
`(74) Attorney, Agent, or Firm — Steptoe & Johnson LLP
`
`(57) ABSTRACT
`
`A method and system can generate video content from a
`video. The method and system can include a coordinator, an
`image detector, and an object recognizer. The coordinator
`can be communicatively coupled to a splitter and/or to a
`plurality of demultiplexer nodes. The splitter can be con-
`figured to segment the video. The demultiplexer nodes can
`be configured to extract audio files from the video and/or to
`extract still frame images from the video. The image detec-
`tor can be configured to detect images of objects in the still
`frame images. The object recognizer can be configured to
`compare an image of an object to a fractal. The recognizer
`can be further configured to update the fractal with the
`image. The coordinator can be configured to embed meta-
`data about the object into the video.
`
`11 Claims, 17 Drawing Sheets
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 1 of 31
`
`
`
`
`
`
`
`
`U.S. Patent
`
`Sep. 21, 2021
`
`Sheet 1 of 17
`
`Video |
`110
`e N |
`Audio to text Image to text
`140
`Natural Natural
`language language
`rocessin '
`150 P g processing
`O\ S/
`
`Combine image text
`and audio text
`
`160
`
`Generate
`video text 170
`
`Figure 1
`
`US 11,126,853 B2
`
`120
`130
`IPR2025-00877
`Patent Owner Exhibit 2003
`
`Page 2 of 31
`
`
`
`
`
`
`
`
`U.S. Patent Sep. 21, 2021 Sheet 2 of 17 US 11,126,853 B2
`
`o NN
`‘/‘M N
`(/l'.“ \*« .
`L Server or servers 220 2
`\»:\ e
`o
`
`Network 230
`User equipment 210
`Figure 2
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 3 of 31
`
`
`
`
`
`
`
`
`U.S. Patent
`
`Sep. 21, 2021 Sheet 3 of 17
`
`Video data
`310
`
`Distributed image
`data processing
`320-1
`
`Distributed image
`data processing
`320-2
`
`4
`
`“\ //
`
`Combine 330
`
`Figure 3
`
`US 11,126,853 B2
`
`Distributed image
`data processing
`320-N
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 4 of 31
`
`
`
`
`
`
`
`
`U.S. Patent
`
`Sep. 21, 2021 Sheet 4 of 17
`
`Audio data
`410
`
`PN
`
`Distributed audio
`data processing
`420-1
`
`Distributed audio
`data processing
`420-2
`
`US 11,126,853 B2
`
`Distributed audio
`data processing
`420-N
`
`N //
`
`Combine 430
`
`Figure 4
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 5 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`Sheet 5 of 17
`
`Sep. 21, 2021
`
`U.S. Patent
`
`uspmeewfioy .
`aoplaun s ,
`
`snAgasy
`WIRNPEAY 0%
`
`sosi
`
`dopy weia ueptdionay wfiwficmmwm
`doy preesg womeBousy
`
`rhag by @ piagg
`
`doy 883 yees Mnma._aamw - ey OB
`
`NG Aseucon o B uofioy pey 1../
`
`LDRERBVICH T
`, - ey)
`
`P {sonpans {ponpuss F el BREOR
`
`gy Appunesy Apquessy feie § UOIDBIAQ 10 ay pdy uoRRnBALnN § -1 Oftd OIDRY
`BEPQ IRVALFRITS FUBRCY s 1
`
`)
`
`Bil-4 CAEA
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 6 of 31
`
`
`
`
`
`
`
`
`U.S. Patent Sep. 21, 2021 Sheet 6 of 17 US 11,126,853 B2
`
`Figure 6
`
`3
`Vigdeo Puss Audiz Data
`¥ ¥
`
`/ Widiex:
`
`Aoy
`I Isspatf TRt InpudfOutet
`¢
`¢
`[TV
`
`)
`
`Dapriutad imsge
`
`Distribured sudic
`Proeussen FRCOSSRS
`Segsusnied &
`et i Fuld Vi
`Tramarig F
`Sendts d
`Voringder MLE
`Povsess :
`; v
`Frugre Topics — Teek Topics
`3 gcw.w ,,.,w'“‘g
`"““WW H H i
`g § g - P
`TR M topics BT
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 7 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`Sheet 7 of 17
`
`Sep. 21, 2021
`
`U.S. Patent
`
`L Eswfi
`
`I sanscnn i iy s
`
`Grooag unnbauriy Gai
`
`s =
`
`WEBAS Whi SRS
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 8 of 31
`
`
`
`
`
`
`
`
`U.S. Patent Sep. 21, 2021 Sheet 8 of 17 US 11,126,853 B2
`
`Figure &
`
`uotMmag afeuy uopusioroy afeuy
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 9 of 31
`
`
`
`
`
`
`
`
`U.S. Patent Sep. 21, 2021 Sheet 9 of 17 US 11,126,853 B2
`
`Figure 9
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 10 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`Sheet 10 of 17
`
`Sep. 21, 2021
`
`U.S. Patent
`
`ws 01 231
`
`Sujuies
`
`291080488y
`
`SWay
`p18007
`{23084
`
`uonudosay
`
`swawdag 5 J0IBUIPI00) ¢ xnuiag
`
`elpan
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 11 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`Sheet 11 of 17
`
`Sep. 21, 2021
`
`U.S. Patent
`
`B e N
`
`.,
`
`{ Bumssooud-aig pug “m.._@.~
`3, 3
`% S
`
`P
`
`%
`
`]
`
`] 1 omn31g
`
`Spoy sy rRueg
`o SuSUERT
`sEsu PuspDY
`EOURT JNBURIDDTD
`‘Fasieay i
`
`=1 BuEE a0
`FOIEG IR TR S0
`S JExEph Y TauEg
`
`Sissanaug
`S0 FRPOE
`BE-C 6t pita N gl
`o uswsat
`BT SRINGUIEID
`JEIBUIDITET
`
`Sl i
`i=smy Sipsy RO
`
`sEEEu
`HiZE ool yusuuses
`mipEpp ey
`Spoh JoERdYTRUS
`
`] o .(L/(
`
`w..i.....
`
`Ky
`
`",
`
`x.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\*
`
`A,
`
`oy wEun .
`Jusilsg g o
`
`)
`
`=
`
`Bumzanodd
`[PUED SUUTIRG
`o} TRV 8T .
`IBEEY ERENEUY
`Brovitiakin vy
`
`RN SRIEIERRON
`JEmEpdi e
`
`ERnhey JuEsennig |
`cEnUSnIEY
`JOIBHDIDET
`
`ettt ettt o
`
`y
`
`A
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 12 of 31
`
`
`
`
`
`
`
`
`U.S. Patent
`
`v \,\
`o -,
`rd e,
`7 ierieres
`&Y Training
`~
`S
`
`Rodhules
`
`Sep. 21, 2021
`
`Sheet 12 of 17
`
`idzrtify Processing
`Windidas
`
`forimesge E
`
`Figure 13
`
`US 11,126,853 B2
`
`s P fseen Batraction
`: o mage
`
`For Fossible
`Dretmetion
`Tompsre
`Sgairst
`Fefersnoe
`
`SeersReccgniton
`
`Parcentags
`
`&
`
`Zrore g
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 13 of 31
`
`
`
`
`
`
`
`
`U.S. Patent Sep. 21, 2021 Sheet 13 of 17 US 11,126,853 B2
`
`§ Mear Frame %
`"{m:;: ing Coamip Eeffi::gj'
`
`Spstem identifies
`Fossible 3d Matches
`
`.-f""" a %""‘fi.
`P
`
`~,
`",
`77777777777 Irmzge Data Fegzed
`
`Y
`", Semre For Srsbpsiz
`
`”””””””””””””””””””””””””””” 0 3
`f.x-*" e
`e o,
`
`0 2 Borated
`
`~High Dorfidencs "*a«.}_&
`
`bngse Uets sddsd to
`
`o ’ A
`o, o
`v
`
`Criginal Medis Azsst
`RaPrzoessed
`
`Figure 14
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 14 of 31
`
`
`
`
`
`
`
`
`U.S. Patent
`
`o
`W
`
`=
`
`{ biadia dsser %
`
`", . m"’l
`
`RN NRARANRNRARANANS
`
`Sep. 21, 2021
`
`Embadded
`\‘”'4
`
`", ‘-"
`
`A TypaSupports ™
`
`Sheet 14 of 17
`
`US 11,126,853 B2
`
`Lreats Subtitie {3RT
`
`stz Btream
`
`Craate VP Filgwith
`dared Mets-Uzim
`
`.................... —
`
`.................... —
`
`Ernbad ¥0AF bate-
`
`.
`
`Erbet Wata-Dara
`
`{opy Filetn
`
`vt Szt Dopy
`
`Fesuiteet Folder
`
`A,
`
`N
`:x.\.\.\.\.\.\.\.\.fi Eont S]
`S, ¢
`
`T, o
`
`Figure 15
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 15 of 31
`
`
`
`
`
`
`
`
`U.S. Patent Sep. 21, 2021 Sheet 15 of 17 US 11,126,853 B2
`
`Figure 16
`
`A
`H
`
`G of 20
`
`&
`Y
`H
`
`IPR2025-00877
`Patent Owner Exhibit 2003
`Page 16 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`Sheet 16 of 17
`
`et
`
`[ N3y
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 17 of 31
`
`Sep. 21, 2021
`
`U.S. Patent
`
`
`
`
`
`
`
`
`
`U.S. Patent Sep. 21, 2021 Sheet 17 of 17 US 11,126,853 B2
`
`> o
`oo
`&
`-
`-]
`T 4]
`>
`IPR2025-00877
`Patent Owner Exhibit 2003
`
`Page 18 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`1
`VIDEO TO DATA
`
`CLAIM OF PRIORITY
`
`This application is a continuation of U.S. application Ser.
`No. 15/197,727, filed Jun. 29, 2016, now U.S. Pat. No.
`10,204,274, which is incorporated by reference in its
`entirety.
`
`TECHNICAL FIELD
`
`The present invention relates to a method and a system for
`generating various and useful data from source media, such
`as videos and other digital content. The data can be embed-
`ded within the source media or combined with the source
`media for creating an augmented video containing additional
`contextual information.
`
`BACKGROUND
`
`In the field of image contextualization, distributed reverse
`image similarity searching can be used to identify images
`similar to a target image. Reverse image searching can find
`exactly matching images as well as flipped, cropped, and
`altered versions of the target image. Distributed reverse
`image similarity searching can be used to identify symbolic
`similarity within images. Audio-to-text algorithms can be
`used to transcribe text from audio. An exemplary application
`is note-taking software. Audio-to-text, however, lacks
`semantic and contextual language understanding.
`
`SUMMARY
`
`The present invention is generally directed to a method to
`generate data from video content, such as text and/or image-
`related information. A server executing the method can be
`directed by a program stored on a non-transitory computer-
`readable medium. The video text can be, for example, a
`context description of the video.
`
`An aspect can include a system for generating data from
`a video. The system can include a coordinator, an image
`detector, and an object recognizer. The coordinator can be
`communicatively coupled to a splitter and/or to a plurality of
`demultiplexer nodes. The splitter can be configured to
`segment the video. The demultiplexer nodes can be config-
`ured to extract audio files from the video and/or to extract
`still frame images from the video. The image detector can be
`configured to detect images of objects in the still frame
`images. The object recognizer can be configured to compare
`an image of an object to a fractal. The recognizer can be
`further configured to update the fractal with the image. The
`coordinator can be configured to embed metadata about the
`object into the video.
`
`In some embodiments, the metadata can include a time-
`stamp and/or a coordinate location of the object in one or
`more of the still frame images. The coordinator can be
`configured to create additional demultiplexer processing
`capacity. The coordinator can be configured to create addi-
`tional demultiplexer nodes, e.g., when the demultiplexer
`nodes reach at least 80% of processing capacity.
`
`In other embodiments, the demultiplexer nodes can gen-
`erate a confidence score based on a comparison of the image
`and the fractal. In yet other embodiments, the recognizer can
`generate a confidence score based on a comparison of the
`image and the fractal.
`
`Another aspect can include a method to generate data
`from a video. The method can include segmenting the video
`
`10
`
`15
`
`20
`
`25
`
`30
`
`35
`
`40
`
`45
`
`50
`
`55
`
`60
`
`65
`
`2
`
`into video segments, extracting an audio file from a segment
`of the video segments, extracting a video frame file of still
`frames from the segment, detecting an image of an object in
`the still frames, recognizing the object as a specific object,
`updating an object-specific fractal with the image, and
`embedding metadata in the video about the specific object.
`
`In some embodiments, the metadata can include a time-
`stamp and/or a coordinate location of the object in one or
`more of the still frames. The metadata can include a recog-
`nition confidence score. The method can further include
`distributing the video segments across a plurality of proces-
`sors. The method can include extracting a plurality of video
`frame files, such as all of the video segments, by a plurality
`of parallel processors.
`
`In other embodiments, the video can be a stereoscopic
`three-dimensional video.
`
`In yet other embodiments, the method can include gen-
`erating text based on extracted audio file and/or applying
`natural language processing to the text. The method can
`include determining context associated with the video based
`on the natural language processing.
`
`In some embodiments, the method can include processing
`the video frame file to extract image text. The object can be
`a face or a logo. The object can be recognized as a three-
`dimensional rotation of a known object.
`
`In other embodiments, a three-dimensional fractal can be
`updated, e.g., with the image of the object. The method can
`include generating a content-rich video based on the video
`and the metadata.
`
`Another aspect can include a system for generating data
`from a video. The system can include a coordinator, an
`image detector, and an object recognizer. The coordinator
`can be communicatively coupled to a splitter and/or to a
`plurality of demultiplexer nodes. The splitter can be con-
`figured to segment the video. The demultiplexer nodes can
`be configured to extract audio files from the video and/or to
`extract still frame images from the video. The image detec-
`tor can be configured to detect images of objects in the still
`frame images. The object recognizer can be configured to
`compare an object image of an object to a fractal. The
`recognizer can be further configured to update the fractal
`with the object image. The coordinator can be configured to
`generate one or more metadata streams corresponding to the
`images. The one or more metadata streams can include
`timestamps corresponding to the images. The coordinator
`can be configured to embed the metadata streams in the
`video.
`
`In some embodiments, the metadata streams can be
`embedded in the video as subtitle resource tracks.
`
`In other embodiments, the system can be accessible over
`a network via application program interfaces (APIs).
`
`In yet other embodiments, the coordinator can be further
`configured to output the video according to multiple video
`formats. For example, the coordinator can be configured to
`automatically generate data files in a variety of formats for
`delivery independent of the video. The system in some
`embodiments can embed data as a stream, as a wrapper,
`and/or as a subtitle resource track. The coordinator can be
`configured to read/write to/from Media Asset Management
`Systems, Digital Asset Management Systems, and/or Con-
`tent Management Systems.
`
`In some embodiments, the system can be configured to
`capture the geolocation of objects in a video. The system can
`be configured to derive a confidence score for each instance
`of recognition. The system can be configured to apply
`natural language processing, for example, for associative
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 19 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`3
`
`terms and/or to apply contextual analysis of corresponding
`data points (such as audio, objects, etc.) to verify accuracy.
`
`An aspect can include a method of creating data from a
`video by machine recognition. The method can include
`extracting an audio file from the video, segmenting the video
`into video frames of still images, distributing the video
`segments to N processors, wherein N is an integer greater
`than one, generating a timestamped transcript from the audio
`file, associating the timestamped transcript with correspond-
`ing video frames, deriving topics from the audio file based
`on natural language processing, recognizing an object from
`still images, using a reference database to identify the object,
`and embedding, within the video, data based on a recognized
`object, the topics, and the timestamped transcript.
`
`In some embodiments, the video can be a virtual reality
`video file or a traditional video vile. Data based on the
`recognized object can include a geolocation.
`
`In other embodiments, the method can include generating
`a plurality of video files. Each of the video files can include
`the video and the embedded data. Each of the plurality of
`video files can be generated in a different format.
`
`In other embodiments, the method can include generating
`a confidence score. The score can be associated with the
`recognized object. The method can include analyzing the
`still images to determine context of the video.
`
`DESCRIPTION OF THE DRAWINGS
`
`The present invention is further described in the detailed
`description which follows, in reference to the noted plurality
`of drawings by way of non-limiting examples of certain
`embodiments of the present invention, in which like numer-
`als represent like elements throughout the several views of
`the drawings, and wherein:
`
`FIG. 1 illustrates an exemplary workflow in certain
`embodiments.
`
`FIG. 2 illustrates an embodiment of image data process-
`ing.
`
`FIG. 3 illustrates aspects of image data processing.
`
`FIG. 4 illustrates aspects of audio data processing.
`
`FIG. 5 illustrates various exemplary aspects of embodi-
`ments of the present invention.
`
`FIG. 6 illustrates a flow diagram of a present embodiment
`FIG. 7 illustrates exemplary architecture of a present
`embodiment.
`
`FIG. 8 illustrates a flow diagram of an embodiment of
`image recognition.
`
`FIG. 9 illustrates an embodiment of a graphical user
`interface of the present invention.
`
`FIG. 10 illustrates exemplary system architecture with an
`exemplary process flow.
`
`FIG. 11 illustrates an exemplary process for distributed
`demultiplexing and preparation of source media files.
`
`FIG. 12 illustrates exemplary distributed processing and
`aggregation.
`
`FIG. 13 illustrates an exemplary process for improved
`recognition based on near frame proximity.
`
`FIG. 14 illustrates an exemplary process for improved
`recognition based on partial three-dimensional matching.
`
`FIG. 15 illustrates an exemplary process for embedding
`extracted data to original source files as metadata.
`
`FIG. 16 depicts an exemplary interface showing a 360°
`image from a virtual reality video file and embedded meta-
`data.
`
`FIG. 17 is an image of the Kress Building in Ft. Worth
`Tex. as taken by a fisheye lens, as used in virtual reality
`images.
`
`25
`
`35
`
`40
`
`45
`
`55
`
`4
`
`FIG. 18 depicts a distorted image after calibration accord-
`ing to present embodiments.
`
`DETAILED DESCRIPTION
`
`A detailed explanation of the system and method accord-
`ing to exemplary embodiments of the present invention are
`described below. Exemplary embodiments described,
`shown, and/or disclosed herein are not intended to limit the
`claims, but rather, are intended to instruct one of ordinary
`skill in the art as to various aspects of the invention. Other
`embodiments can be practiced and/or implemented without
`departing from the scope and spirit of the claimed invention.
`
`The present invention is generally directed to system,
`device, and method of generating data from source media,
`such as images, video, and audio. Video can include two-
`dimensional video and/or stereoscopic three-dimensional
`video such as virtual reality (VR) files. The generated data
`can include text and information relating to context, sym-
`bols, brands, features, objects, faces and/or topics found in
`the source media. In an embodiment, the video-to-data
`engine can perform the functions directed by programs
`stored in a computer-readable medium. That is, the embodi-
`ments can include hardware (such as circuits, processors,
`memory, user and/or hardware interfaces, etc.) and/or soft-
`ware (such as computer-program products that include com-
`puter-useable instructions embodied on one or more com-
`puter-readable media).
`
`The various video-to-data techniques, methods, and sys-
`tems described herein can be implemented in part or in
`whole using computer-based systems and methods. Addi-
`tionally, computer-based systems and methods can be used
`to augment or enhance the functionality described herein,
`increase the speed at which the functions can be performed,
`and provide additional features and aspects as a part of, or
`in addition to, those described elsewhere herein.
`
`Various computer-based systems, methods, and imple-
`mentations in accordance with the described technology are
`presented below.
`
`A video-to-data engine can be embodied by a computer or
`a server and can have an internal or external memory for
`storing data and programs such as an operating system (e.g.,
`DOS, Windows2000™, Windows XP™, Windows NT™,
`0S/2, UNIX, Linux, Xbox OS, Orbis OS, and FreeBSD)
`and/or one or more application programs. The video-to-data
`engine can be implemented by a computer or a server
`through tools of a particular software development kit
`(SDK). Examples of application programs include computer
`programs implementing the techniques described herein for
`lyric and multimedia customization, authoring applications
`(e.g., word processing programs, database programs, spread-
`sheet programs, or graphics programs) capable of generating
`documents, files, or other electronic content; client applica-
`tions (e.g., an Internet Service Provider (ISP) client, an
`e-mail client, or an instant messaging (IM) client) capable of
`communicating with other computer users, accessing vari-
`ous computer resources, and viewing, creating, or otherwise
`manipulating electronic content; and browser applications
`(e.g., Microsoft’s Internet Explorer) capable of rendering
`standard Internet content and other content formatted
`according to standard protocols such as the Hypertext Trans-
`fer Protocol (HTTP). One or more of the application pro-
`grams can be installed on the internal or external storage of
`the computer. Application programs can be externally stored
`in or performed by one or more device(s) external to the
`computer.
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 20 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`5
`
`The computer or server can include a central processing
`unit (CPU) for executing instructions in response to com-
`mands, and a communication device for sending and receiv-
`ing data. One example of the communication device can be
`a modem. Other examples include a transceiver, a commu-
`nication card, a satellite dish, an antenna, a network adapter,
`or some other mechanism capable of transmitting and
`receiving data over a communications link through a wired
`or wireless data pathway.
`
`The computer or server can also include an input/output
`interface that enables wired or wireless connection to vari-
`ous peripheral devices. In one implementation, a processor-
`based system of the computer can include a main memory,
`preferably random access memory (RAM), and can also
`include a secondary memory, which can be a tangible
`computer-readable medium. The tangible computer-read-
`able medium memory can include, for example, a hard disk
`drive or a removable storage drive, a flash based storage
`system or solid-state drive, a floppy disk drive, a magnetic
`tape drive, an optical disk drive (Blu-Ray, DVD, CD drive),
`magnetic tape, paper tape, punched cards, standalone RAM
`disks, lomega Zip drive, etc. The removable storage drive
`can read from or write to a removable storage medium. A
`removable storage medium can include a floppy disk, mag-
`netic tape, optical disk (Blu-Ray disc, DVD, CD) a memory
`card (CompactFlash card, Secure Digital card, Memory
`Stick), paper data storage (punched card, punched tape), etc.,
`which can be removed from the storage drive used to
`perform read and write operations. As will be appreciated,
`the removable storage medium can include computer soft-
`ware or data.
`
`In alternative embodiments, the tangible computer-read-
`able medium memory can include other similar means for
`allowing computer programs or other instructions to be
`loaded into a computer system. Such means can include, for
`example, a removable storage unit and an interface.
`Examples of such can include a program cartridge and
`cartridge interface (such as found in video game devices), a
`removable memory chip (such as an EPROM or flash
`memory) and associated socket, and other removable stor-
`age units and interfaces, which allow software and data to be
`transferred from the removable storage unit to the computer
`system.
`
`An embodiment of video-to-data engine operation is
`illustrated in FIG. 1. At 110, a video stream is presented. The
`video stream can be in one or more of the formats (but not
`limited to): Advanced Video Codec High Definition
`(AVCHD), Audio Video Interlaced (AVI), Flash Video For-
`mat (FLV), Motion Picture Experts Group (MPEG), Win-
`dows Media Video (WMV), or Apple QuickTime (MOV),
`h.264 (MP4).
`
`The engine can extract audio data and image data (e.g.
`images or frames forming the video) from the video stream.
`The engine can detect and identify objects, faces, logos, text,
`music, sounds and spoken language in video by means of
`demultiplexing and extracting features from the video and
`passing those features into a distributed system as the video
`loads into the network I/O buffer. In some embodiments, the
`video stream and the extracted image data can be stored in
`a memory or storage device such as those discussed above.
`A copy of the extracted image data can be used for process-
`ing.
`
`The system and method can include dialog extraction.
`Language and vocabulary models can be included in the
`system to support desired languages. Multiple languages can
`be incorporated into the system and method. The engine can
`process audio media containing multiple audio tracks as
`
`25
`
`40
`
`45
`
`50
`
`6
`
`separate tracks or as a single track. Text extraction can be
`optimized by utilizing audio segments of various lengths in
`time. For example, if a segment of audio is greater than one
`minute, the engine can split the audio track in half. In this
`case, the engine can first analyze that specific sequence for
`dialog at the timestamp of the potential split. If the segment
`at the split contains audio, the system can split the audio at
`the next silent point in the track to avoid splitting tracks
`mid-word. Each segment is processed using the Kaldi pro-
`cess for speech recognition and dialog extraction. Segments
`can be subsequently processed through, for example, LIUM
`speaker diarization. Results can be applied to a result
`datastore for analysis or later processing.
`
`An example of the image data processing is illustrated in
`FIG. 3. The video-to-data engine can segment the video into
`chunks for distributed, or parallel, processing as shown
`schematically in FIG. 3. Distributed processing in this
`context can mean that the processing time for analyzing a
`video from beginning to end is a fraction of the play time of
`the video. This can be accomplished by breaking the pro-
`cesses into sections and processing them simultaneously.
`The images and audio can each be broken up into pieces
`such that the meaning of a continuous message is preserved.
`At 120, the video-to-data engine performs an image data
`processing on the video stream. In FIG. 3, the image data
`310 can be segmented into N segments and processed in
`parallel (e.g., distributed processing 320-1 to 320-N), allow-
`ing for near real-time processing.
`
`An example of the video image data processing can be
`symbol (or object) based. Using an image processing tech-
`nique such as color edge detection, a symbol of a screen or
`an image of the video can be isolated. The symbol can be
`identified using an object template database. For example,
`the symbol includes 4 legs and a tail, and when matched with
`the object template database, the symbol may be identified
`as a dog. The object template database can be adaptive and
`therefore, the performance would improve with usage.
`
`Other image data processing techniques can include
`image extraction, high-level vision and symbol detection,
`figure-ground separation, depth and motion perception.
`These and/or other image data processing techniques can be
`utilized to build a catalogue and/or repository of extracted
`objects. Recognized information about the extracted
`objects—such as object type, context, brands/logos, vehicle
`type/make/model, clothing worn, celebrity name, etc.—can
`be added to a file associated with the extracted object and/or
`used to augment metadata in the video from which the image
`data was processed.
`
`Another example of video image processing can be color
`segmentation. The colors of an image (e.g., a screen) of the
`video can be segmented or grouped. The result can be
`compared to a database using color similarity matching.
`
`Based on the identified symbol, a plurality of instances of
`the symbol can be compared to a topic database to identify
`a topic (such as an event). For example, the result may
`identify the dog (symbol) as running or jumping. The topic
`database can be adaptive to improve its performance with
`usage.
`
`Thus, using the processing example above, text describing
`a symbol of the video and topic relating to the symbol can
`be generated, as is illustrated in FIG. 9. Data generated from
`an image and/or from audio transcription can be time
`stamped, for example, according to when it appeared, was
`heard, and/or according to the video frame from which it
`was pulled. The time-stamped data can be physically asso-
`ciated with the video as metadata embedded at the relevant
`portion of video.
`
`IPR2025-00877
`
`Patent Owner Exhibit 2003
`
`Page 21 of 31
`
`
`
`
`
`
`
`
`US 11,126,853 B2
`
`7
`
`At 330, the engine combines the topics as an array of keys
`and values with respect to the segments. The engine can
`segment the topics over a period of time and weight the
`strength of each topic. Further, the engine applies the topical
`metadata to the original full video. The image topics can be
`stored as topics for the entire video or each image segment.
`The topic generation process can be repeated for all identi-
`fiable symbols in a video in a distributed process. The
`outcome would be several topical descriptors of the content
`within a video. An example of the aggregate information that
`can be derived using the above example would be a deter-
`mination that the video presented a dog, which was jumping,
`on the beach, with people, by a resort.
`
`Although further described herein, image detection can be
`considered a process of determining if a pattern or patterns
`exist in an image and whether the pattern or patterns meet
`criteria of a face, image, and/or text. If the result is positive,
`image recognition can be employed. Image recognition can
`generally be considered matching of detected objection to
`known objects and/or matching through machine learning.
`Generally speaking, detection and recognition, while shar-
`ing various aspects, are distinct.
`
`Identifying various objects in an image can be a difficult
`task. For example, locating or segmenting and positively
`identifying an object in a given frame or image can yield
`false positives—Ilocating but wrongfully identifying an
`object. Therefore, present embodiments can be utilized to
`eliminate or reduce false positives, for example, by using
`context. As one example, if the audio soundtrack of a video
`is an announcer calling a football game, then identification
`of ball in a given frame as a basketball can be assigned a
`reduced probability or weighting. As another example of
`using context, if a given series of image frames from a video
`is positively or strongly identified as a horse race, then
`identifying an object to be a mule or donkey can be given a
`reduced weight.
`
`Using the context or arrangement of certain objects in a
`given still or static image to aid in computer visual recog-
`nition accuracy can be an extremely difficult task given
`certain challenges associated with partially visible or self-
`occluded objects, lack of objects, and/or faces, and/or words
`or an overly cluttered image, etc. However, the linear
`sequencing of frames from a video—as opposed to a stand-
`alone image—avails itselfto a set images {images x-y} from
`which context can be derived. This contextual methodology
`can be viewed as systematic detection of probable image
`false positives by identifying an object from one video frame
`(or image) as an anomaly when compared to and associated
`with a series of image frames both prior and subsequent to
`the purported anomaly. According to the objects, faces,
`words, etc. of a given set of frames (however defined), a
`probability can be associated with an identified anomaly to
`determine whether an image is a false positive and, if so,
`what other likely results can be.
`
`In certain instances, identification of an individual can be
`a difficult task. For example, facial recognition can become
`difficult when an individual’s face is obstructed by another
`object like a football, a baseball helmet, a musical instru-
`ment, or other obstructions. An advantage of some embodi-
`ments described herein can include the ability to identify an
`individual without identification of the individual’s face.
`Embodiments can use contextual information such as asso-
`ciations of objects, text, and/or other context within an
`image or video. As one example, a football player scores a
`touchdown but rather than identifying the player using facial
`recognition, the player can be identified by object recogni-
`tio



