nerdexam
Google

GENERATIVE-AI-LEADER · Question #45

A travel app asks users to take a photo of a famous landmark and then returns a written overview with historical notes and nearby attractions. The system's capability to interpret the picture and…

The correct answer is B. A multimodal learning model. This scenario requires understanding visual content from a photo and then generating a textual explanation. That means the system consumes one modality as an image and produces another modality as text. This cross-modality capability is exactly what a multimodal approach…

Multimodal Generative AI

Question

A travel app asks users to take a photo of a famous landmark and then returns a written overview with historical notes and nearby attractions. The system's capability to interpret the picture and produce natural language output reflects what kind of model?

Options

  • AAn image classification model
  • BA multimodal learning model
  • CA time-series forecasting model
  • DA text-only unimodal model

How the community answered

(31 responses)
  • B
    90% (28)
  • C
    6% (2)
  • D
    3% (1)

Explanation

This scenario requires understanding visual content from a photo and then generating a textual explanation. That means the system consumes one modality as an image and produces another modality as text. This cross-modality capability is exactly what a multimodal approach provides, since it jointly handles vision and language to produce coherent natural language output based on visual input.

Topics

#Multimodal AI#Generative Models#Computer Vision#Natural Language Generation

Community Discussion

No community discussion yet for this question.

Full GENERATIVE-AI-LEADER Practice