Enhancing Insights into Traffic Accident Risk with Multimodal Retrieval-Augmented Generation
マルチモーダルRAGを用いた交通事故リスク推定
- Delivery
- Available on this site
- Format
- Price
- Non-members (tax incl.):¥1,100 Members (tax incl.):¥880
- Publication code
- 20264497
- Paper/Info type
- Journal of Society of Automotive Engineers of Japan
Vol.80 No.7
- Pages
- 126-131(Total 6 p)
- Date of publication
- Jul 2026
- Publisher
- JSAE
- Language
- Japanese
Detailed Information
| Category(J) | ホットトピックス Translation |
|---|---|
| Category(E) | Hot Topics |
| Author(J) | 1) 伊藤 修, 2) 千葉 大幹 |
| Author(E) | 1) Osamu Ito, 2) Motoki Chiba |
| Affiliation(J) | 1) 本田技研工業, 2) 本田技研工業 |
| Abstract(J) | 視覚言語モデル(Vision-Language Model; VLM)を交通事故リスク推定タスクへ適応させるためのファインチューニングを効率化するアノテーション手法を提案する。少量のラベル付きデータを用い、マルチモーダルRAG (Retrieval-Augmented Generation) を活用することで未見の画像に対して事故リスクを生成するというものであり、本手法により推定性能が向上することを確認した。 Translation |
| Abstract(E) | We propose an efficient annotation method for fine-tuning Vision-Language Models (VLMs) specifically for traffic accident risk estimation. By leveraging multimodal Retrieval-Augmented Generation (RAG) with a limited set of labeled data, our approach generates accurate risk descriptions for unseen images. Experimental results demonstrate that this method effectively enhances the model’s risk estimation performance, providing a scalable solution for scene understanding without the need for extensive manual labeling. |