[HTML][HTML] Summary of chatgpt-related research and perspective towards the future of large language models
This paper presents a comprehensive survey of ChatGPT-related (GPT-3.5 and GPT-4)
research, state-of-the-art large language models (LLM) from the GPT series, and their …
research, state-of-the-art large language models (LLM) from the GPT series, and their …
Tools, techniques, datasets and application areas for object detection in an image: a review
J Kaur, W Singh - Multimedia Tools and Applications, 2022 - Springer
Object detection is one of the most fundamental and challenging tasks to locate objects in
images and videos. Over the past, it has gained much attention to do more research on …
images and videos. Over the past, it has gained much attention to do more research on …
Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models
Large Vision-Language Models (LVLMs) have recently played a dominant role in
multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation …
multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation …
Ocr-free document understanding transformer
Understanding document images (eg, invoices) is a core but challenging task since it
requires complex functions such as reading text and a holistic understanding of the …
requires complex functions such as reading text and a holistic understanding of the …
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Vision Language Models (VLMs), which extend Large Language Models (LLM) by
incorporating visual understanding capability, have demonstrated significant advancements …
incorporating visual understanding capability, have demonstrated significant advancements …
Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Pre-training of text and layout has proved effective in a variety of visually-rich document
understanding tasks due to its effective model architecture and the advantage of large-scale …
understanding tasks due to its effective model architecture and the advantage of large-scale …
Docformer: End-to-end transformer for document understanding
We present DocFormer-a multi-modal transformer based architecture for the task of Visual
Document Understanding (VDU). VDU is a challenging problem which aims to understand …
Document Understanding (VDU). VDU is a challenging problem which aims to understand …
On the hidden mystery of ocr in large multimodal models
Large models have recently played a dominant role in natural language processing and
multimodal vision-language learning. However, their effectiveness in text-related visual …
multimodal vision-language learning. However, their effectiveness in text-related visual …
Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding
In this work, we present DocPedia, a novel large multimodal model (LMM) for versatile OCR-
free document understanding, capable of parsing images up to 2560× 2560 resolution …
free document understanding, capable of parsing images up to 2560× 2560 resolution …
Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents
Key information extraction (KIE) from document images requires understanding the
contextual and spatial semantics of texts in two-dimensional (2D) space. Many recent …
contextual and spatial semantics of texts in two-dimensional (2D) space. Many recent …