Select Language

AI社区

公开数据集

Flickr图片数据集,Flickr 图像字幕数据集

Flickr图片数据集,Flickr 图像字幕数据集

8.2G
269 浏览
0 喜欢
0 次下载
0 条讨论
NLP,Image Data,Computer Vision Classification

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30......

数据结构 ? 8.2G

    README.md

    The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains, linking mentions of the same entities across different captions for the same image, and associating them with 276k manually annotated bounding boxes. Such annotations are essential for continued progress in automatic image description and grounded language understanding. They enable us to define a new benchmark for localization of textual entity mentions in an image. We present a strong baseline for this task that combines an image-text embedding, detectors for common objects, a color classifier, and a bias towards selecting larger objects. While our baseline rivals in accuracy more complex state-of-the-art models, we show that its gains cannot be easily parlayed into improvements on such tasks as image-sentence retrieval, thus underlining the limitations of current methods and the need for further research.

    暂无相关内容。
    暂无相关内容。
    • 分享你的想法
    去分享你的想法~~

    全部内容

      欢迎交流分享
      开始分享您的观点和意见,和大家一起交流分享.
    所需积分:35 去赚积分?
    • 269浏览
    • 0下载
    • 0点赞
    • 收藏
    • 分享