マルチモーダルRAGを用いた交通事故リスク推定
Enhancing Insights into Traffic Accident Risk with Multimodal Retrieval-Augmented Generation
- 提供方法
- 本サイト上にてダウンロード・閲覧可
- 形態
- 価格
- 一般価格(税込):¥1,100 会員価格(税込):¥880
- 文献番号
- 20264497
- 文献・情報種別
- 会誌「自動車技術」
Vol.80 No.7
- 掲載ページ
- 126-131(Total 6 p)
- 発行年月
- 2026年 7月
- 出版社
- (公社)自動車技術会
- 言語
- 日本語
書誌事項
| カテゴリ | ホットトピックス |
|---|---|
| カテゴリ(英) | Hot Topics 翻訳 |
| 著者 | 1) 伊藤 修, 2) 千葉 大幹 |
| 著者(英) | 1) Osamu Ito, 2) Motoki Chiba |
| 勤務先 | 1) 本田技研工業, 2) 本田技研工業 |
| 抄録 | 視覚言語モデル(Vision-Language Model; VLM)を交通事故リスク推定タスクへ適応させるためのファインチューニングを効率化するアノテーション手法を提案する。少量のラベル付きデータを用い、マルチモーダルRAG (Retrieval-Augmented Generation) を活用することで未見の画像に対して事故リスクを生成するというものであり、本手法により推定性能が向上することを確認した。 |
| 抄録(英) | We propose an efficient annotation method for fine-tuning Vision-Language Models (VLMs) specifically for traffic accident risk estimation. By leveraging multimodal Retrieval-Augmented Generation (RAG) with a limited set of labeled data, our approach generates accurate risk descriptions for unseen images. Experimental results demonstrate that this method effectively enhances the model’s risk estimation performance, providing a scalable solution for scene understanding without the need for extensive manual labeling. 翻訳 |