ASTRID临床问答评估数据集

包含370条临床问题-答案-上下文三元组,用于评估RAG临床问答系统,涵盖白内障术后随访等多场景。

Ufonia Limited, 约克大学Ufonia Limited, 约克大学
arXiv
2025-01-14 更新
浏览 20
文本白内障手术随访临床问答

基本信息

模态
文本
创建/更新时间
2025-01-14

资源简介

ASTRID数据集由Ufonia Limited和约克大学创建,用于评估基于检索增强生成(RAG)的临床问答系统。包含来自真实患者的白内障手术后随访问题,以及急诊、临床和非临床领域的问题,共370条问题-答案-上下文三元组,分为FaithfulnessQAC(238条)和UniqueQAC(132条)两个子集。该数据集旨在解决现有评估指标在临床和对话场景中的不足,确保生成的回答在临床上是准确且有用的。

原始链接

http://arxiv.org/abs/2501.08208v1

访问原始数据

官方服务

如需原始数据获取支持或标注服务,请联系我们。

帮我联系

下载信息

注册下载

Tips: 该数据集需要在对应的数据源网站注册通过后,才能进行数据下载,注册有对应要求,或者需要收费。

暂未开放

公开下载

Tips: 该数据集属于公开下载,应该可以免费公开下载。

免登录

有偿下载

Tips: 该数据集 Qianfanghub 可以协助提供有偿下载服务,注意,服务不针对数据相关产权,只是技术服务费。

提供高速下载与技术交付服务(收技术服务费,非数据销售)

暂未开放

千方医数集,医疗数据集部分,是为社区服务的公开医疗数据集搜索引擎,并不存储或者下载原始的任何数据。 如果您有其他医疗数据需求,可以和客服联系,或者下工单。我们有强大的三甲医疗机构帮助您提供个性化的医疗数据定制、采集、标注服务。

使用方式

数据集说明

ASTRID临床问答评估数据集 对应论文数据集(arXiv 预印本)。

数据获取指引

  1. 打开论文页面获取作者与项目信息:https://arxiv.org/abs/2501.08208v1
  2. 论文 Data Availability / Code Availability 章节标注了数据实际托管位置;
  3. 获取到实际数据链接后,按对应平台标准方式下载。

论文摘要:Abstract:Large Language Models (LLMs) have shown impressive potential in clinical question answering (QA), with Retrieval Augmented Generation (RAG) emerging as a leading approach for ensuring the factual accuracy of model responses. However, current automated RAG metrics perform poorly in clinical and conversational use cases. Using clinical human evaluations of responses is expensive, unscalable, and not conducive to the continuous iterative development of RAG systems. To address these challenges, we introduce ASTRID - an Automated and Scalable TRIaD for evaluating clinical QA systems leveraging RAG - consisting of three metrics: Context Relevance (CR), Refusal Accuracy (RA), and Conversational Faithfulness (CF). Our novel evaluation metric, CF, is designed to better capture the faithfulness of a model's response to the knowledge base without penalising conversational elements. To validate our triad, we curate a dataset of over 200 real-world patient questions posed to an LLM-based QA agent during surgical follow-up for cataract surgery - the highest volume operation in the world - augmented with clinician-selected questions for emergency, clinical, and non-clinical out-of-domain scenarios. We demonstrate that CF can predict human ratings of faithfulness better than existing definitions for conversational use cases. Furthermore, we show that evaluation using our triad consisting of CF, RA, and CR exhibits alignment with clinician assessment for inappropriate, harmful, or unhelpful responses. Finally, using nine different LLMs, we demonstrate that the three metrics can

论文页面:https://arxiv.org/abs/2501.08208v1

精度瓶颈?数据缺失?

当前公开数据无法满足您的算法精度?千方提供针对 白内障 的高质量、多模态真实临床数据定制解决方案。

获取专属数据定制方案