You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: content/english/events/registration.md
+19-3Lines changed: 19 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -74,14 +74,30 @@ quick_navigation:
74
74
75
75
{{< accordions_area_open id="event" >}}
76
76
77
-
{{< accordion_item_open title="Masterclass 'Testing General Purpose AI (GPAI) applications'" id="event" background_color="#ffffff" tag1="3 November 2026" tag2="masterclass" tag3="in-person" image="/images/events/20260625_GPAI_event.png" >}}
77
+
{{< accordion_item_open title="Masterclass 'Testing General Purpose AI (GPAI) applications'" id="event" background_color="#ffffff" tag1="3 November 2026" tag2="masterclass" tag3="in-person" image="/images/events/20261103_GPAI_event.png" >}}
78
78
79
79
{{< promo_bar index="0" >}}
80
80
81
81
<br>
82
82
83
83
#### Description
84
-
The capabilities of GPAI models and systems continue to improve. GPAI systems now match or exceed expert-level performance on capabilities such as “reasoning”, “problem solving” and “graduate-level skills”. Yet it is often unclear what these concepts actually capture. In this masterclass, attendees gain state-of-the-art insights about the latest developments in the rapidly advancing field of GPAI benchmarking.
84
+
As GPAI models continue to improve, systems are increasingly touted as matching or exceeding expert performance in “problem solving”, “scientific reasoning” or “real-world software engineering.” These claims are based on common tests for different capabilities, known as “benchmarks”.
85
+
86
+
Yet, there is often misconception of what benchmarks actually measure, what conclusions about a system we can actually make and how to evaluate what actually matters for a use case. Just as other technology-driven industries, like healthcare and aviation, AI benchmarking needs reliable standards. With best practices in their infancy, it can be daunting for professionals to navigate GPAI evaluation.
87
+
88
+
In this masterclass, we distill and make accessible the most valuable insights from the field of GPAI benchmarking. The course covers:
89
+
- GPAI under the AI Act and benchmarking for systems with systemic risk
90
+
- Prominent industry benchmarks and adaptations to GPT-NL
91
+
- Insights from Algorithm Audit’s own benchmarking work
92
+
- Current issues in the field from a scientific perspective
93
+
- The design aspects that determine the quality of a benchmark
94
+
95
+
After attending, participants will be able to:
96
+
- Understand relevant obligations for GPAI systems under the AI Act
97
+
- Grasp how benchmarks tend to work “under the hood”
98
+
- Navigate common evaluation repositories and documentation
99
+
- Assess the suitability of existing benchmarks for their own work
100
+
- Reason when a custom approach is needed and how it may look
85
101
86
102
#### Date
87
103
3 November 2026
@@ -106,7 +122,7 @@ The Hague Conference Centre (New Babylon), Anna van Buerenplein 29, 2595 DA Den
106
122
#### Audience
107
123
Professionals from private and public sector who regularly work with GPAI applications, such as implementation of generative AI solutions in work processes, testing GPAI capabilities and/or working on AI policy.
Copy file name to clipboardExpand all lines: content/nederlands/events/registration.md
+20-3Lines changed: 20 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -74,14 +74,31 @@ quick_navigation:
74
74
75
75
{{< accordions_area_open id="event" >}}
76
76
77
-
{{< accordion_item_open title="Masterclass 'Testen van General Purpose AI (GPAI) toepassingen'" id="event" background_color="#ffffff" tag1="3 november 2026" tag2="masterclass" tag3="op locatie" image="/images/events/20260625_GPAI_event.png" >}}
77
+
{{< accordion_item_open title="Masterclass 'Testen van General Purpose AI (GPAI) toepassingen'" id="event" background_color="#ffffff" tag1="3 november 2026" tag2="masterclass" tag3="op locatie" image="/images/events/20261103_GPAI_event.png" >}}
78
78
79
79
{{< promo_bar index="0" >}}
80
80
81
81
<br>
82
82
83
83
#### Beschrijving
84
-
De capaciteiten van GPAI-modellen en -systemen blijven zich ontwikkelen. GPAI-systemen evenaren of overtreffen inmiddels expertniveau op vaardigheden als “redeneren”, “probleemoplossen” en “academische vaardigheden”. Toch is vaak onduidelijk wat deze begrippen daadwerkelijk omvatten. In deze masterclass krijgen deelnemers state-of-the-art inzichten in de nieuwste ontwikkelingen op het snel evoluerende gebied van GPAI-benchmarking.
84
+
Naarmate GPAI-modellen steeds beter worden, wordt er steeds vaker beweerd dat systemen de prestaties van experts evenaren of overtreffen op het gebied van ‘probleemoplossing’, ‘wetenschappelijk redeneren’ of ‘softwareontwikkeling in de praktijk’. Deze beweringen zijn gebaseerd op gangbare tests voor verschillende vaardigheden, ook wel ‘benchmarks’ genoemd.
85
+
86
+
Toch bestaat er vaak een misvatting over wat benchmarks nu eigenlijk meten, welke conclusies we daadwerkelijk over een systeem kunnen trekken en hoe we kunnen beoordelen wat er voor een specifieke toepassing echt toe doet. Net als in andere technologiegedreven sectoren, zoals de gezondheidszorg en de luchtvaart, heeft AI-benchmarking betrouwbare standaarden nodig. Omdat best practices nog in de kinderschoenen staan, kan het voor professionals best lastig zijn om hun weg te vinden in GPAI-evaluaties.
87
+
88
+
In deze masterclass distilleren we de meest waardevolle inzichten uit het gebied van GPAI-benchmarking en maken we die toegankelijk. De cursus behandelt:
89
+
- GPAI volgens de AI-wet en benchmarking voor systemen met systeemrisico’s
90
+
- Toonaangevende benchmarks uit de sector en toepassing op GPT-NL
91
+
- Inzichten uit het eigen benchmarkwerk van Algorithm Audit
92
+
- Actuele kwesties in het vakgebied vanuit een wetenschappelijk perspectief
93
+
- De ontwerpaspecten die de kwaliteit van een benchmark bepalen
94
+
95
+
Na het volgen van de training kun je:
96
+
- De relevante verplichtingen voor GPAI-systemen onder de AI-wet begrijpen
97
+
- Begrijpen hoe benchmarks ‘achter de schermen’ werken
98
+
- Je weg vinden in gangbare evaluatiedatabases en documentatie
99
+
- Beoordelen of bestaande benchmarks geschikt zijn voor je eigen werk
100
+
- Bepalen wanneer een aangepaste aanpak nodig is en hoe die eruit zou kunnen zien
101
+
85
102
86
103
#### Datum
87
104
3 november 2026
@@ -106,7 +123,7 @@ The Hague Conference Centre (New Babylon), Anna van Buerenplein 29, 2595 DA Den
106
123
#### Doelgroep
107
124
Professionals uit de private en publieke sector die regelmatig werken met GPAI-toepassingen, zoals het implementeren van generatieve AI-oplossingen in werkprocessen, het testen van GPAI-capaciteiten en/of het werken aan AI-beleid.
0 commit comments