Amazon released Marengo 3.0 on an unspecified date, saying it enables natural language search through video, audio, and image content. It is the company's first multimodal embedding update since the introduction of Amazon Bedrock Knowledge Bases.

Amazon reported Marengo 3.0 generates 512-dimensional vectors for video, audio, images, and text, measured on a compact, storage-efficient embedding space. That compares with earlier models that used less efficient vector spaces.

Marengo 3.0 is built on Amazon Bedrock Knowledge Bases and targets industries such as media, sports analytics, and security. Availability begins with general availability, initially for users in the US East (N. Virginia) and US West (N. California) Regions.

"Video and media assets remain largely unsearchable by meaning," said Amazon. "Building semantic search over video today requires stitching together a complex pipeline of transcription services, frame extraction pipelines, embedding models, vector databases, and synchronization logic."

The announcement follows the launch of Amazon Bedrock Knowledge Bases. Amazon framed the significance as unlocking the full value of media assets through natural language search without complex infrastructure.

Amazon did not say how Marengo 3.0 compares to other embedding models, and raised the open question of how the model handles large video files. The company said users can start integrating the model through the Amazon Bedrock Retrieve API.

Source: awsml