`
`a2 United States Patent
`
`ao) Patent No.: US 10,218,954 B2
`
`Lakhani et al. 45) Date of Patent: *Feb. 26, 2019
`(54) VIDEO TO DATA (52) US. CL
`] CPC ............ HO4N 9/8715 (2013.01); HO4N 9/80
`(71) Applicant: CELLULAR SOUTH, INC., (2013.01); HO4N 9/802 (2013.01); GIOL
`Ridgeland, MS (US) 15/26 (2013.01)
`. . . (58) Field of Classification Search
`(72) - Inventors: g:fff;f&gfii“é’mciifiy}l%n @%}SOH CPC ... GO1L 5/18; GO1L, 5/26; GO1L, 5/265; G10L,
`= 3 e . . .
`MS (US); Allison A. Talley, Ridgeland, 25/48; GI0L 25/51; G10L 25/54;
`MS (US) (Continued)
`(73) Assignee: EEESEEAI\I}[ Ssg}g{H’ INC., (56) References Cited
`] ] o ] U.S. PATENT DOCUMENTS
`(*) Notice: Subject to any disclaimer, the term of this
`patent is extended or adjusted under 35 7,536,413 B1* 5/2009 Mohan ............... GOG6F 17/3071
`U.S.C. 154(b) by 0 days. 9,615,136 B1* 4/2017 Emery .......... HO4N 21/47202
`(b) by 0 day .
`This patent is subject to a terminal dis- (Continued)
`claimer.
`OTHER PUBLICATIONS
`(21) Appl. No.: 14/910,698 o ) ) ) o
`Notification Concerning Transmittal of International Preliminary
`(22) PCT Filed: Feb. 7, 2015 Report on Patentability (Chapter I of the Patent Cooperation Treaty)
`dated Aug. 18, 2016, issued in International Application No. PCT/
`(86) PCT No.: PCT/US2015/014940 US2015/014940.
`§ 371 (c)(1), (Continued)
`(2) Date: Feb. 6, 2016
`Primary Examiner — Paras D Shah
`(87) PCT Pub. No.: 'WO2015/120351 (74) Attorney, Agent, or Firm — Steptoe & Johnson LLP
`PCT Pub. Date: Aug. 13, 2015
`57 ABSTRACT
`(65) Prior Publication Data 7
`A method and system can generate video content from a
`US 2016/0358632 Al Dec. 8, 2016 video. The method and system can include generating audio
`Related U.S. Application Data files and image files from the video, distributing the audio
`(63) Continuation of application No. 14/175,741, filed on files and the image files across a plurality of processors and
`Feb. 7, 2014, now Pat. No. 9,940,972. processing the audio files and the image files in parallel. The
`(Continued) audio files associated with the video to text and the image
`files associated with the video to video content can be
`(51) Int.CL converted. The text and the video content can be cross-
`GI10L 15/00 (2013.01) referenced with the video.
`GI0L 25/00 (2013.01)
`(Continued) 18 Claims, 9 Drawing Sheets
`
`Video
`
`10
`
`A
`
`S
`
`Audio to text
`
`140
`
`Image to text
`
`120
`
`I
`
`I
`
`Natural
`language
`processing
`
`150
`
`Natural
`language
`processing 130
`
`N
`
`J
`
`Combine image text
`and audio text
`
`160
`
`I
`
`Generate
`video text 70
`
`IPR2025-00877
`
`Patent Owner Exhibit 2002
`
`Page 1 of 19
`
`
`
`
`
`
`
`
`US 10,218,954 B2
`
`Page 2
`Related U.S. Application Data 2007/0106685 A1* 5/2007 Houh .............. GOG6F 17/30796
`o o 2007/0112630 Al* 5/2007 GO6K 9/72
`(60) Provisional application No. 62/021,666, filed on Jul. 705/14.1
`7,2014, provisional application No. 61/866,175, filed 2008/0120646 A1* 5/2008 Stern .................. G06Q 30/02
`on Aug. 15, 2013. 725/34
`2008/0232696 Al* 9/2008 Kasahara .......... GO6K 9/00664
`382/224
`(1) IG110t.6FCl}7/30 (2006.01) 2008/0276266 Al* 11/2008 Huchital ............... G06Q 30/02
`: 725/32
`Go6T 7/00 (2017.01) 2009/0006191 A1* 1/2009 Arankalle .............. GO6Q 30/02
`HO4N 9/87 (2006.01) 705/14.71
`HO4N 9/802 (2006.01) 2009/0006375 Al* 12009 Lax ..o HO4N 21/435
`HO4IN 9/80 (2006.01) 2009/0157407 Al 6/2009 Yamabe et al.
`GIOL 15/26 (2006.01) 2009/0199235 Al* 82009 Surendran .............. G06Q 30/02
`. I ) 725/34
`(58) Field of Classification Search 2011/0069936 Al 3/2011 Johnson et al.
`CPC ... G10L 25/57, GO6F 17/30, GOG6F 17/30058; 2011/0292992 Al* 12/2011 Sirivara ........... HO4N 21/23418
`GO6F 17/3074; GO6F 17/30781; GO6T 375/240.01
`7/0081; GO6K 2209/01 2012/0254917 Al* 10/2012 Burkitt .............. GO6F 17/30817
`. 725/40
`USPC ... 7047231, 523}127760?7%72212/%121/;0221?15/7631, 2013/0174045 Al 7/2013 Sarukkai et al.
`. 2 ’ L0 2014/0314391 Al* 102014 Kim ....cccovvnveenneen. Gl11B 27/11
`See application file for complete search history. 386/248
`2015/0019206 Al* 1/2015 Wilder .............. GO6K 9/00302
`(56) References Cited 704/9
`U.S. PATENT DOCUMENTS
`OTHER PUBLICATIONS
`9,652,477 B2* 5/2017 Balinsky ............. GOG6F 17/3028
`2001/0020954 Al 9/2001 Hull et al. International Search Report dated Jun. 3, 2015, issued in Interna-
`2003/0112261 Al1* 6/2003 Zhang ................ G11B 27/22 tional Application No. PCT/US2015/014940.
`N 715/716 Written Opinion of the International Searching Authority dated Jun.
`2003/0112265 Al 6/2003 Zhang ... GOGF 1;/135%55 3, 2015, issued in International Application No. PCT/US2015/
`014940.
`3k
`2005/0108004 Al 3/2005 0N v G107Loi/52/(7)§ Deng, Yining, and B.S. Manjunath, “Content-based search of video
`2006/0179453 Al* 82006 Kadie ..oooovoeovonn GO6Q 30/02 using color, texture, and motion.” Image Processing, 1997, retrieved
`725/34 from: http:/vision.ece.uc.sb.edu/publications/97ICIPCBS pdf.
`2006/0212897 ALl* 9/2006 Li .cccooovviniininennn HO4H 60/58
`725/32 * cited by examiner
`
`IPR2025-00877
`
`Patent Owner Exhibit 2002
`
`Page 2 of 19
`
`
`
`
`
`
`
`
`U.S. Patent
`
`140
`
`150
`
`Feb. 26, 2019
`
`Sheet 1 of 9
`
`Video
`
`7 N
`
`Audio to text
`
`10
`
`I
`
`Image to text
`
`Natural
`language
`processing
`
`|
`
`Natural
`language
`processing
`
`N/
`
`Combine image text
`and audio text
`
`160
`
`J
`
`Generate
`video text
`
`170
`
`FIG. 1
`
`US 10,218,954 B2
`
`120
`130
`IPR2025-00877
`Patent Owner Exhibit 2002
`
`Page 3 of 19
`
`
`
`
`
`
`
`
`U.S. Patent Feb. 26, 2019 Sheet 2 of 9 US 10,218,954 B2
`
`{ Video data
`| 310
`
`|
`
`[
`
`Distributed image Distributed image Distributed image
`data processing data processing es e data processing
`320-1 320-2 320-N
`
`Combine 330
`FIG. 2
`
`IPR2025-00877
`Patent Owner Exhibit 2002
`Page 4 of 19
`
`
`
`
`
`
`
`
`U.S. Patent
`
`Feb. 26, 2019 Sheet 3 of 9
`
`Audio data
`410
`
`US 10,218,954 B2
`
`AL NN
`
`Distributed audio
`data processing
`420-1
`
`Distributed audio
`data processing
`420-2
`
`Combine 430
`
`FIG. 3
`
`Distributed audio
`data processing
`420-N
`
`NN\ /S
`
`IPR2025-00877
`
`Patent Owner Exhibit 2002
`
`Page 5 of 19
`
`
`
`
`
`
`
`
`U.S. Patent
`
`Feb. 26, 2019 Sheet 4 of 9 US 10,218,954 B2
`
`h
`
`( Network 230
`
`N
`Va
`
`User equipment 210
`
`FIG. 4
`
`IPR2025-00877
`
`Patent Owner Exhibit 2002
`
`Page 6 of 19
`
`
`
`
`
`
`
`
`U.S. Patent Feb. 26, 2019 Sheet 5 of 9 US 10,218,954 B2
`
`2
`
`2
`
`£
`
`s
`
`PSRN
`
`g
`
`FIG. 5
`
`43
`
`Lrretteitiatit e tin e
`g
`%
`H
`
`—
`e
`A
`e
`
`i
`
`e
`NPPERES
`BPRERE S
`ARSI
`
`ot
`e
`
`e
`
`[RESRRRRRNNER
`
`IPR2025-00877
`Patent Owner Exhibit 2002
`Page 7 of 19
`
`
`
`
`
`
`
`
`U.S. Patent Feb. 26, 2019 Sheet 6 of 9 US 10,218,954 B2
`
`Figure 6
`
`R
`
`Audie sty
`
`Distriusnd Aulls ©
`g
`
`eshrRusted &
`
`ProrRmes
`
`“favt Topins
`
`IPR2025-00877
`Patent Owner Exhibit 2002
`Page 8 of 19
`
`
`
`
`
`
`
`
`US 10,218,954 B2
`
`Sheet 7 of 9
`
`Feb. 26, 2019
`
`U.S. Patent
`
`|
`
`S6L
`
`S
`NN
`
`‘.
`
`06/
`
`&
`
`¥ \\sx\\\\x\\x‘.\xvx“&
`g s oioe
`Y
`
`.
`%
`
`(s
`
`R
`
`oy
`
`251
`
`iy
`
`&3
`
`SRR
`
`S
`
`SSRGS
`
`Y
`
`%,
`U ip ot porisnd
`
`eaavansansit
`
`IPR2025-00877
`
`Patent Owner Exhibit 2002
`
`Page 9 of 19
`
`
`
`
`
`
`
`
`U.S. Patent Feb. 26, 2019 Sheet 8 of 9 US 10,218,954 B2
`
`800
`810
`
`IPR2025-00877
`Patent Owner Exhibit 2002
`Page 10 of 19
`
`
`
`
`
`
`
`
`U.S. Patent Feb. 26, 2019 Sheet 9 of 9 US 10,218,954 B2
`
`FIG. 9
`
`IPR2025-00877
`Patent Owner Exhibit 2002
`Page 11 of 19
`
`
`
`
`
`
`
`
`US 10,218,954 B2
`
`1
`VIDEO TO DATA
`
`CLAIM OF PRIORITY
`
`This application claims the benefit under 35 USC 371 to
`International Application No. PCT/US2015/014940, filed
`Feb. 7, 2015, which claims priority to U.S. patent applica-
`tion Ser. No. 14/175,741, filed Feb. 7, 2014, which claims
`priority to U.S. Provisional Patent Application No. 61/866,
`175, filed on Aug. 15, 2013, and claims priority to U.S.
`Provisional Patent Application No. 62/021,666, filed Jul. 7,
`2014, each of which is incorporated by reference in its
`entirety.
`
`TECHNICAL FIELD
`
`The present invention relates to a method and a system for
`generating various and useful data from videos.
`
`BACKGROUND
`
`In the field of image contextualization, distributed reverse
`image similarity searching can be used to identify images
`similar to a target image. Reverse image searching can find
`exactly matching images as well as flipped, cropped, and
`altered versions of the target image. Distributed reverse
`image similarity searching can be used to identify symbolic
`similarity within images. Audio-to-text algorithms can be
`used to transcribe text from audio. An exemplary application
`is note-taking software. Audio-to-text, however, lacks
`semantic and contextual language understanding.
`
`SUMMARY
`
`The present invention is generally directed to a method to
`generate data from video content, such as text and/or image-
`related information. A server executing the method can be
`directed by a program stored on a non-transitory computer-
`readable medium. The video text can be, for example, a
`context description of the video.
`
`An aspect of the method can include generating text from
`an image of the video, converting audio associated with the
`video to text, extracting topics from the text converted from
`the audio, cross-referencing the text generated from the
`image of the video and the topics extracted from audio
`associated with the video, and generating video text based
`on a result of the cross-referencing.
`
`In some embodiments, natural language processing can be
`applied to the generation of text from an image of the video,
`converting audio associated with the video to text, or both.
`
`In other embodiments, the text from the image of the
`video can be generated by identifying context, a symbol, a
`brand, a feature, an object, and/or a topic in the image of the
`video.
`
`In yet other embodiments, the text from the image can be
`generated by first segmenting images of the video, and then
`converting the segments of images to text in parallel. The
`text from the audio can be generated by first segmenting
`images of the audio, and then converting the segments of
`images to text in parallel. The audio can be segmented at
`spectrum thresholds. The generated text may be of different
`sizes. The size of the text can be adjusted by a ranking or
`scoring function that, for example, can adjust the text size
`based on confidence in the description or relevance to a
`search inquiry. The text can describe themes, identification
`of objects or other information of interest.
`
`10
`
`15
`
`20
`
`25
`
`30
`
`35
`
`40
`
`45
`
`50
`
`55
`
`60
`
`65
`
`2
`
`In some embodiments, the method can include generating
`advertising and/or product or service recommendations
`based on video content. The video content can be text,
`context, symbols, brands, features, objects, and/or topics
`related to or found in the video. An advertisement and/or
`product or service recommendations can be placed at a
`specific time in the video based on the video content and/or
`section symbol of a video image. The advertisement and/or
`product or service recommendations can also be placed at a
`specific time as part of the video player, e.g., side panel, and
`also may be placed on a second screen. In some embodi-
`ments, the method can include directing when one or more
`advertisements can be placed in a predetermined context at
`a preferred time.
`
`BRIEF DESCRIPTION OF THE DRAWINGS
`
`The present invention is further described in the detailed
`description which follows, in reference to the noted plurality
`of drawings by way of non-limiting examples of certain
`embodiments of the present invention, in which like numer-
`als represent like elements throughout the several views of
`the drawings, and wherein:
`
`FIG. 1 illustrates an embodiment of present invention.
`
`FIG. 2 illustrates an embodiment of image data process-
`ing.
`
`FIG. 3 illustrates an embodiment of audio data process-
`ing.
`
`FIG. 4 illustrates another embodiment of present inven-
`tion.
`
`FIG. 5 illustrates various exemplary embodiments of
`present invention.
`
`FIG. 6 illustrates a flow diagram of an embodiment
`
`FIG. 7 illustrates an embodiment of the architecture of the
`present invention.
`
`FIG. 8 illustrates a flow diagram of an embodiment of
`image recognition.
`
`FIG. 9 illustrates an embodiment of a graphical user
`interface of the present invention.
`
`DETAILED DESCRIPTION
`
`A detailed explanation of the system and method accord-
`ing to exemplary embodiments of the present invention are
`described below. Exemplary embodiments described,
`shown, and/or disclosed herein are not intended to limit the
`claims, but rather, are intended to instruct one of ordinary
`skill in the art as to various aspects of the invention. Other
`embodiments can be practiced and/or implemented without
`departing from the scope and spirit of the claimed invention.
`
`The present invention is generally directed to system,
`device, and method of generating content from video files,
`such as text and information relating to context, symbols,
`brands, features, objects, faces and/or topics found in the
`images of such videos. In an embodiment, the video-to-
`content engine can perform the functions directed by pro-
`grams stored in a computer-readable medium. That is, the
`embodiments may take the form of a hardware embodiment
`(including circuits), a software embodiment, or an embodi-
`ment combining software and hardware. The present inven-
`tion can take the form of a computer-program product that
`includes computer-useable instructions embodied on one or
`more computer-readable media.
`
`The various video-to-content techniques, methods, and
`systems described herein can be implemented in part or in
`whole using computer-based systems and methods. Addi-
`
`tionally, computer-based systems and methods can be used IPR2025-00877
`Patent Owner Exhibit 2002
`
`Page 12 of 19
`
`
`
`
`
`
`
`
`US 10,218,954 B2
`
`3
`
`to augment or enhance the functionality described herein,
`increase the speed at which the functions can be performed,
`and provide additional features and aspects as a part of or in
`addition to those described elsewhere in this document.
`Various computer-based systems, methods and implemen-
`tations in accordance with the described technology are
`presented below.
`
`A video-to-content engine can be embodied by the a
`general-purpose computer or a server and can have an
`internal or external memory for storing data and programs
`such as an operating system (e.g., DOS, Windows 2000™,
`Windows XP™, Windows NT™, OS/2, UNIX or Linux)
`and one or more application programs. Examples of appli-
`cation programs include computer programs implementing
`the techniques described herein for lyric and multimedia
`customization, authoring applications (e.g., word processing
`programs, database programs, spreadsheet programs, or
`graphics programs) capable of generating documents or
`other electronic content; client applications (e.g., an Internet
`Service Provider (ISP) client, an e-mail client, or an instant
`messaging (IM) client) capable of communicating with other
`computer users, accessing various computer resources, and
`viewing, creating, or otherwise manipulating electronic con-
`tent; and browser applications (e.g., Microsoft’s Internet
`Explorer) capable of rendering standard Internet content and
`other content formatted according to standard protocols such
`as the Hypertext Transfer Protocol (HTTP). One or more of
`the application programs can be installed on the internal or
`external storage of the general-purpose computer. Alterna-
`tively, application programs can be externally stored in or
`performed by one or more device(s) external to the general-
`purpose computer.
`
`The general-purpose computer or server may include a
`central processing unit (CPU) for executing instructions in
`response to commands, and a communication device for
`sending and receiving data. One example of the communi-
`cation device can be a modem. Other examples include a
`transceiver, a communication card, a satellite dish, an
`antenna, a network adapter, or some other mechanism
`capable of transmitting and receiving data over a commu-
`nications link through a wired or wireless data pathway.
`
`The general-purpose computer or server may also include
`an input/output interface that enables wired or wireless
`connection to various peripheral devices. In one implemen-
`tation, a processor-based system of the general-purpose
`computer can include a main memory, preferably random
`access memory (RAM), and can also include a secondary
`memory, which may be a tangible computer-readable
`medium. The tangible computer-readable medium memory
`can include, for example, a hard disk drive or a removable
`storage drive, a flash based storage system or solid-state
`drive, a floppy disk drive, a magnetic tape drive, an optical
`disk drive (Blu-Ray, DVD, CD drive), magnetic tape, paper
`tape, punched cards, standalone RAM disks, lomega Zip
`drive, etc. The removable storage drive can read from or
`write to a removable storage medium. A removable storage
`medium can include a floppy disk, magnetic tape, optical
`disk (Blu-Ray disc, DVD, CD) a memory card (Compact-
`Flash card, Secure Digital card, Memory Stick), paper data
`storage (punched card, punched tape), etc., which can be
`removed from the storage drive used to perform read and
`write operations. As will be appreciated, the removable
`storage medium can include computer software or data.
`
`In alternative embodiments, the tangible computer-read-
`able medium memory can include other similar means for
`allowing computer programs or other instructions to be
`loaded into a computer system. Such means can include, for
`
`20
`
`25
`
`30
`
`40
`
`45
`
`55
`
`4
`
`example, a removable storage unit and an interface.
`Examples of such can include a program cartridge and
`cartridge interface (such as the found in video game
`devices), a removable memory chip (such as an EPROM or
`flash memory) and associated socket, and other removable
`storage units and interfaces, which allow software and data
`to be transferred from the removable storage unit to the
`computer system.
`
`An embodiment of video-to-content engine operation is
`illustrated in FIG. 1. At 110, a video stream is presented. The
`video stream may be in format of (but not limited to):
`Advanced Video Codec High Definition (AVCHD), Audio
`Video Interlaced (AVI), Flash Video Format (FLU), Motion
`Picture Experts Group (MPEG), Windows Media Video
`(WMV), or Apple QuickTime (MOV), h.264 (MP4).
`
`The engine can extract audio data and image data (e.g.
`images or frames forming the video) from the video stream.
`
`In some embodiments, the video stream and the extracted
`image data can be stored in a memory or storage device such
`as those discussed above. A copy of the extracted image data
`can be used for processing.
`
`At 120, the video-to-content engine performs an image
`data processing on the video stream. An example of the
`image data processing is illustrated in FIG. 2. In FIG. 2, the
`image data 310 can be segmented into N segments and
`processed in parallel (e.g., distributed processing 320-1 to
`320-N), allowing for near real-time processing.
`
`An example of the video image data processing can be
`symbol (or object) based. Using image processing technique
`such as color edge detection, a symbol of a screen or an
`image of the video can be isolated. The symbol can be
`identified using an object template database. For example,
`the symbol includes 4 legs and a tail, and when matched with
`the object template database, the symbol may be identified
`as a dog. The object template database can be adaptive and
`therefore, the performance would improve with usage.
`
`Other image data processing techniques may include
`image extraction, high-level vision and symbol detection,
`figure-ground separation, depth and motion perception.
`
`Another example of video image processing can be color
`segmentation. The colors of an image (e.g., a screen) of the
`video can be segmented or grouped. The result can be
`compared to a database using color similarity matching.
`
`Based on the identified symbol, a plurality of instances of
`the symbol can be compared to a topic database to identify
`a topic (such as an event). For example, the result may
`identify the dog (symbol) as running or jumping. The topic
`database can be adaptive to improve its performance with
`usage.
`
`Thus, using the processing example above, text describing
`a symbol of the video and topic relating to the symbol may
`be generated, as is illustrated in FIG. 9. Data generated from
`an image and/or from audio transcription can be time
`stamped, for example, according to when it appeared, was
`heard, and/or according to the video frame from which it
`was pulled.
`
`At 330, the engine combines the topics as an array of keys
`and values with respect to the segments. The engine can
`segment the topics over a period of time and weight the
`strength of each topic. Further, the engine applies the topical
`meta-data to the original full video. The image topics can be
`stored as topics for the entire video or each image segment.
`The topic generation process can be repeated for all identi-
`fiable symbols in a video in a distributed process. The
`outcome would be several topical descriptors of the content
`within a video. An example of the aggregate information that
`would be derived using the above example would be under-
`
`IPR2025-00877
`
`Patent Owner Exhibit 2002
`
`Page 13 of 19
`
`
`
`
`
`
`
`
`US 10,218,954 B2
`
`5
`
`standing that the video presented a dog, which was jumping,
`on the beach, with people, by a resort.
`
`Identifying various objects in an image can be a difficult
`task. For example, locating (segmenting) and positively
`identifying an object in a given frame or image can yield
`false positives—Ilocating but wrongfully identifying an
`object. Therefore, present embodiments can be utilized to
`eliminate false positives, for example, by using context. As
`one example, if the audio soundtrack of a video is an
`
`6
`
`and assigning probabilities. For example, neuro-linguistic
`programming (NLP), neural network programming, or deep
`neural networks can be utilized to achieve sufficient nar-
`rowing and weighting. For further example, based on a
`contextual review of a large number of objects over a period
`of time, a series of nodes in parallel and/or in series can be
`developed by the processor. Upon initial recognition of
`objects and context, these nodes can assign probabilities to
`the initial identification of the objection with each node in
`
`announcer calling a football game, then identification of ball 10 . . text and further description t th
`in a given frame as a basketball can be assigned a reduced H,L 1:)511ng COE x anf v b.e r EE;CII: ption ho dnla row the
`probability or weighting. As another example of using probabi istic choices Qrano JeCt'. .t er methodo ogles can
`context, if a given series of image frames from a video is be u.tlhzed to determine and/or utilize context as described
`positively or strongly identified as a horse race, then iden- herein. . . .
`tifying an object to be a mule or donkey can be given a 15 | Namral language processing can be useful in .Creatmg an
`reduced weight. intuitive and/or user-friendly computer-human interaction.
`Using the context or arrangement of certain objects in a In some embodiments, the system can select semantics or
`given still or static image to aid in computer visual recog- topics, following certain rules, from a plurality of possible
`nition accuracy can be an extremely difficult task given semantics or topics, can give them weight based on strength
`certain challenges associated with partially visible or self- 20 of context, and/or can do this a distributed environment. The
`occluded objects, lack of objects, and/or faces, and/or words natural language processing can be augmented and/or
`or an overly cluttered image, etc. However, the linear improved by implementing machine-learning. A large train-
`sequencing of frames from a video—as opposed to a stand- ing set of data can be obtained from proprietary or publicly
`alone image—avails itselfto a set images {images x-y} from available resources. For example, CBS News maintains a
`which context can be derived. This contextual methodology 25 database of segments and episodes of “60-Minutes” with full
`can be viewed as systematic detection of probable image transcripts, which can be useful for building a training set
`false positives by identifying an object from one video frame and for unattended verification of audio segmentation. The
`(or image) as an anomaly when compared to and associated machine learning can include ensemble learning based on
`with a series of image frames both prior and subsequent to the concatenation of several classifiers, i.e. cascade classi-
`the purported anomaly. According to the objects, faces, 30 fiers.
`words, etc. of a given set of frames (however defined), a At 130, an optional step of natural language processing
`probability can be associated with an identified anomaly to can be applied to the image text. For example, based on
`determine whether an image may be a false positive and, if dictionary, grammar, and a knowledge database, the text
`so, what other likely results should be. extracted from video images can be modified as the video-
`In certain instances, identification of an individual can be 35 to-content engine selects primary semantics from a plurality
`a difficult task. For example, facial recognition can become of possible semantics. In some embodiments, the system and
`difficult when an individual’s face is obstructed by another method can incorporate a Fourier transform of the audio
`object like a football, a baseball helmet, a musical instru- signal. Such filtering can improve silence recognition, which
`ment, or other obstructions. An advantage of some embodi- can be useful for determining proper placement of commas
`ments described herein can include the ability to identify an 40 and periods in the text file.
`individual without identification of the individual’s face. In parallel, at 140, the video-to-content engine can per-
`Embodiments can use contextual information such as asso- form audio-to-text processing on audio data associated with
`ciations of objects, text, and/or other context within an the video. For example, for a movie video, the associated
`image or video. As one example, a football player scores a audio may be the dialog or even background music.
`touchdown but rather than identifying the player using facial 45 Inaddition to filtering of the audio signal, images from the
`recognition, the player can be identified by object recogni- video signal can be processed to address, for example, the
`tion of, for example, the player’s team’s logo, text recog- problem of object noise in a given frame or image. Often
`nition of the player’s jersey number, and by cross referenc- images are segmented only to locate and positively identify
`ing this data with that team’s roster (as oppose to another one or very few main images in the foreground of a given
`team, which is an example of why the logo recognition can 50 frame. The non-primary or background images are often
`be important). Such embodiments can further learn to iden- treated as noise. Nevertheless, these can provide useful
`tify that player more readily and save his image as data. information, context and/or branding for two examples. To
`Similarly, the audio transcript of a video can be used to fine-tune the amount of object noise cluttering a data set, it
`derive certain context helpful in identifying and correcting can be useful to provide a user with an option to dial image
`or eliminating image false positives. In this way, an image 55 detection sensitivity. For certain specific embodiments,
`anomaly or anomalies identified in a given video frame(s) identification of only certain clearly identifiable faces or
`are associated with time (time stamped) and correlated with large unobstructed objects or band logos can be required
`a time range from the transcribed audio to establish certain with all other image noise disregarded or filtered, which can
`probabilities of accuracy. require less computational processing and image database
`Moreover, the aforementioned methodologies—establish- 60 referencing, in turn reducing costs. However, it may become
`ing probabilities of accuracy of image identification from a necessary or desirable to detect more detail from a frame or
`set of frames and from the audio transcription—can be set of frames. In such circumstances, the computational
`combined to improve the results. thresholds for identification of an object, face, etc. can be
`In some embodiments, a similar context methodology can altered according to a then stated need or desire for non-
`be used to identify unknown objects in a given image by 65 primary, background, obstructed and/or grainy type images.
`
`narrowing a large, or practically infinite, number of possi-
`bilities to a relatively small number of object possibilities
`
`Such image identification threshold adjustment capability
`
`can be implemented, for example, as user-controlled inter- IPR2025-00877
`
`Patent Owner Exhibit 2002
`Page 14 of 19
`
`
`
`
`
`
`
`
`US 10,218,954 B2
`
`7
`
`face, dial, slider, or button, which enables the user to make
`adjustments to suit specific needs or preferences.
`
`An example of the audio data processing is illustrated in
`FIG. 3. In FIG. 3, the audio data 410 can be segmented into
`N segments and processed in parallel (e.g., distributed
`processing 420-1 to 420-N), allowing for near real-time
`processing.
`
`In some embodiments, the segmentation can be per-
`formed by a fixed period of time. In another example, quiet
`periods in the audio data can be detected, and the segmen-
`tation can be defined by the quiet periods. For example, the
`audio data can be processed and converted into a spectrum.
`Locations where the spectrum volatility is below a threshold
`can be detected and segmented. Such locations can represent
`silence or low audio activities in the audio data. The quiet
`periods in the audio data can be ignored, and the processing
`requirements thereof can be reduced.
`
`Audio data and/or segments of audio data can be stored in,
`for example, memory or storage device discussed above.
`Copies of the audio segments can be sent to audio process-
`ing.
`
`The audio data for each segment can be translated into
`text in parallel, for example through distributed computing,
`which can reduce processing time. Various audio analysis
`tools and processes can be used, such as audio feature
`detection and extraction, audio indexing, hashing and
`searching, semantic analysis, and synthesis.
`
`At 430, text for a plurality of segments can then be
`combined. The combination can result in segmented tran-
`scripts and/or a full transcript of the audio data. In an
`embodiment, the topics in each segment can be extracted.
`When combined, the topics in each segment can be given a
`different weight.
`
`The audio topics can be stored as topics for the entire
`video or each audio segment.
`
`At 150, an optional step of natural language processing
`can be applied to the text. For example, based on dictionary,
`grammar, and/or a knowledge database, the text extract from
`the audio stream of a video can be given context, an applied
`sentiment, and topical weightings.
`
`At 160, the topics generated from an image or a frame and
`the topics extracted from audio can be combined. The text
`can be cross-referenced, and topics common to both texts
`would be given additional weights. At 170, the video-to-
`content engine generates video text, such as text describing
`the content of the video, using the result of the combined
`texts and cross reference. For example, key words indicating
`topic and semantic that appear in both texts can be selected
`or emphasized. The output can also include metadata that
`can be time-stamped with frame references. The metadata
`can include the number of frames, the range of frames,
`and/or timestamp references.
`
`FIG. 4 illustrates another embodiment of the present
`invention. User equipment (UE) 210 can communicate with
`a server or servers 220 via a network 230. An exemplary
`embodiment of the system can be implemented over a cloud
`computing network.
`
`For exemplary purposes only, and not to limit one or more
`embodiments herein, FIG. 6 illustrates a flow diagram of an
`embodiment. A video file is first split into video data and
`audio data. A data pipeline, indicated in the figure as Video
`Input/Output, can extract sequences of image frames and can
`warehouse compressed images in a distributed data store as
`image frame data. A distributed computation engine can be
`dedicated to image pre-processing, performing e.g. corner
`and/or edge detection and/or image segmentation. The
`engine can also be dedicated to pattern recognition, e.g. face
`
`5
`
`10
`
`15
`
`20
`
`25
`
`30
`
`35
`
`40
`
`45
`
`50
`
`55
`
`60
`
`8
`
`detection and/or logo recognition, and/or other analysis,
`such as motion tracking Processed data can be sent to one or
`more machines that can combine and/or sort results in a
`time-ordered fashion. Similarly, the Audio Input/Output
`represents a data pipeline for e.g. audio analysis, compres-
`sion, and/or warehousing in a distributed file system. The
`audio can be, for example but not limited to WAV, .MP3, or
`other known formats. Also similarly to the video branch, a
`distributed computation engine c



