GENERATIVE-AI-LEADER · Question #45
A travel app asks users to take a photo of a famous landmark and then returns a written overview with historical notes and nearby attractions. The system's capability to interpret the picture and…
The correct answer is B. A multimodal learning model. This scenario requires understanding visual content from a photo and then generating a textual explanation. That means the system consumes one modality as an image and produces another modality as text. This cross-modality capability is exactly what a multimodal approach…
Question
A travel app asks users to take a photo of a famous landmark and then returns a written overview with historical notes and nearby attractions. The system's capability to interpret the picture and produce natural language output reflects what kind of model?
Options
- AAn image classification model
- BA multimodal learning model
- CA time-series forecasting model
- DA text-only unimodal model
How the community answered
(31 responses)- B90% (28)
- C6% (2)
- D3% (1)
Explanation
This scenario requires understanding visual content from a photo and then generating a textual explanation. That means the system consumes one modality as an image and produces another modality as text. This cross-modality capability is exactly what a multimodal approach provides, since it jointly handles vision and language to produce coherent natural language output based on visual input.
Topics
Community Discussion
No community discussion yet for this question.