Sections
Text Area

RESEARCH THRUSTS

Human-Centric Benchmarking

Image
Image
banner-human
Text Area

 

Benchmarks are standardized measures of performance that allow people to evaluate and compare products. For example, the Massive Multi-task Language Understanding benchmark provides quantified support for the impression that language models have improved by leaps and bounds over the past few years. In the context of artificial intelligence and robotics for geriatric care, benchmarks are essential to ensuring that new technologies are safe and effective.

The ongoing revolution in artificial intelligence poses a challenge to benchmarks. For one thing, many new products have aced established benchmarks, thanks to the surge in performance by language models. This makes it difficult to differentiate competing models and to appreciate incremental improvement. On top of that, many models have been trained on the content of benchmarks, which makes their performance on those benchmarks fragile and unrepresentative of their underlying abilities. As a result, new, hard benchmarks are in demand.

Existing human-centric benchmarks measure artificial intelligence against human performance on standardized tests. Comparing artificial intelligence to a human reference point is an important first step, but it does not yet account for naturalistic interactions with real people. To improve geriatric care effectively, we will develop benchmarks that incorporate real-world scenarios while maintaining rigour. To improve geriatric care safely, we will ensure that these benchmarks are challenging. Any technology involved in caring for older adults must meet a high standard within a well-defined role.