Skip to content

Commit 235ee7a

Browse files
authored
Merge pull request #479 from NGO-Algorithm-Audit/kf_edits
Masterclass description update
2 parents 1b72df1 + b5daeef commit 235ee7a

4 files changed

Lines changed: 39 additions & 6 deletions

File tree

content/english/events/registration.md

Lines changed: 19 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -74,14 +74,30 @@ quick_navigation:
7474

7575
{{< accordions_area_open id="event" >}}
7676

77-
{{< accordion_item_open title="Masterclass 'Testing General Purpose AI (GPAI) applications'" id="event" background_color="#ffffff" tag1="3 November 2026" tag2="masterclass" tag3="in-person" image="/images/events/20260625_GPAI_event.png" >}}
77+
{{< accordion_item_open title="Masterclass 'Testing General Purpose AI (GPAI) applications'" id="event" background_color="#ffffff" tag1="3 November 2026" tag2="masterclass" tag3="in-person" image="/images/events/20261103_GPAI_event.png" >}}
7878

7979
{{< promo_bar index="0" >}}
8080

8181
<br>
8282

8383
#### Description
84-
The capabilities of GPAI models and systems continue to improve. GPAI systems now match or exceed expert-level performance on capabilities such as “reasoning”, “problem solving” and “graduate-level skills”. Yet it is often unclear what these concepts actually capture. In this masterclass, attendees gain state-of-the-art insights about the latest developments in the rapidly advancing field of GPAI benchmarking.
84+
As GPAI models continue to improve, systems are increasingly touted as matching or exceeding expert performance in “problem solving”, “scientific reasoning” or “real-world software engineering.” These claims are based on common tests for different capabilities, known as “benchmarks”.
85+
86+
Yet, there is often misconception of what benchmarks actually measure, what conclusions about a system we can actually make and how to evaluate what actually matters for a use case. Just as other technology-driven industries, like healthcare and aviation, AI benchmarking needs reliable standards. With best practices in their infancy, it can be daunting for professionals to navigate GPAI evaluation.
87+
88+
In this masterclass, we distill and make accessible the most valuable insights from the field of GPAI benchmarking. The course covers:
89+
- GPAI under the AI Act and benchmarking for systems with systemic risk
90+
- Prominent industry benchmarks and adaptations to GPT-NL
91+
- Insights from Algorithm Audit’s own benchmarking work
92+
- Current issues in the field from a scientific perspective
93+
- The design aspects that determine the quality of a benchmark
94+
95+
After attending, participants will be able to:
96+
- Understand relevant obligations for GPAI systems under the AI Act
97+
- Grasp how benchmarks tend to work “under the hood”
98+
- Navigate common evaluation repositories and documentation
99+
- Assess the suitability of existing benchmarks for their own work
100+
- Reason when a custom approach is needed and how it may look
85101

86102
#### Date
87103
3 November 2026
@@ -106,7 +122,7 @@ The Hague Conference Centre (New Babylon), Anna van Buerenplein 29, 2595 DA Den
106122
#### Audience
107123
Professionals from private and public sector who regularly work with GPAI applications, such as implementation of generative AI solutions in work processes, testing GPAI capabilities and/or working on AI policy.
108124

109-
{{< embed_pdf url="/pdf-files/events/activities/20260625_Masterclass_Benchmarking.pdf" width_mobile_pdf="12" width_desktop_pdf="6" >}}
125+
{{< embed_pdf url="/pdf-files/events/activities/20261103_Masterclass_Benchmarking.pdf" width_mobile_pdf="12" width_desktop_pdf="6" >}}
110126

111127
{{< dynamic_form_engine index="0" >}}
112128

content/nederlands/events/registration.md

Lines changed: 20 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -74,14 +74,31 @@ quick_navigation:
7474

7575
{{< accordions_area_open id="event" >}}
7676

77-
{{< accordion_item_open title="Masterclass 'Testen van General Purpose AI (GPAI) toepassingen'" id="event" background_color="#ffffff" tag1="3 november 2026" tag2="masterclass" tag3="op locatie" image="/images/events/20260625_GPAI_event.png" >}}
77+
{{< accordion_item_open title="Masterclass 'Testen van General Purpose AI (GPAI) toepassingen'" id="event" background_color="#ffffff" tag1="3 november 2026" tag2="masterclass" tag3="op locatie" image="/images/events/20261103_GPAI_event.png" >}}
7878

7979
{{< promo_bar index="0" >}}
8080

8181
<br>
8282

8383
#### Beschrijving
84-
De capaciteiten van GPAI-modellen en -systemen blijven zich ontwikkelen. GPAI-systemen evenaren of overtreffen inmiddels expertniveau op vaardigheden als “redeneren”, “probleemoplossen” en “academische vaardigheden”. Toch is vaak onduidelijk wat deze begrippen daadwerkelijk omvatten. In deze masterclass krijgen deelnemers state-of-the-art inzichten in de nieuwste ontwikkelingen op het snel evoluerende gebied van GPAI-benchmarking.
84+
Naarmate GPAI-modellen steeds beter worden, wordt er steeds vaker beweerd dat systemen de prestaties van experts evenaren of overtreffen op het gebied van ‘probleemoplossing’, ‘wetenschappelijk redeneren’ of ‘softwareontwikkeling in de praktijk’. Deze beweringen zijn gebaseerd op gangbare tests voor verschillende vaardigheden, ook wel ‘benchmarks’ genoemd.
85+
86+
Toch bestaat er vaak een misvatting over wat benchmarks nu eigenlijk meten, welke conclusies we daadwerkelijk over een systeem kunnen trekken en hoe we kunnen beoordelen wat er voor een specifieke toepassing echt toe doet. Net als in andere technologiegedreven sectoren, zoals de gezondheidszorg en de luchtvaart, heeft AI-benchmarking betrouwbare standaarden nodig. Omdat best practices nog in de kinderschoenen staan, kan het voor professionals best lastig zijn om hun weg te vinden in GPAI-evaluaties.
87+
88+
In deze masterclass distilleren we de meest waardevolle inzichten uit het gebied van GPAI-benchmarking en maken we die toegankelijk. De cursus behandelt:
89+
- GPAI volgens de AI-wet en benchmarking voor systemen met systeemrisico’s
90+
- Toonaangevende benchmarks uit de sector en toepassing op GPT-NL
91+
- Inzichten uit het eigen benchmarkwerk van Algorithm Audit
92+
- Actuele kwesties in het vakgebied vanuit een wetenschappelijk perspectief
93+
- De ontwerpaspecten die de kwaliteit van een benchmark bepalen
94+
95+
Na het volgen van de training kun je:
96+
- De relevante verplichtingen voor GPAI-systemen onder de AI-wet begrijpen
97+
- Begrijpen hoe benchmarks ‘achter de schermen’ werken
98+
- Je weg vinden in gangbare evaluatiedatabases en documentatie
99+
- Beoordelen of bestaande benchmarks geschikt zijn voor je eigen werk
100+
- Bepalen wanneer een aangepaste aanpak nodig is en hoe die eruit zou kunnen zien
101+
85102

86103
#### Datum
87104
3 november 2026
@@ -106,7 +123,7 @@ The Hague Conference Centre (New Babylon), Anna van Buerenplein 29, 2595 DA Den
106123
#### Doelgroep
107124
Professionals uit de private en publieke sector die regelmatig werken met GPAI-toepassingen, zoals het implementeren van generatieve AI-oplossingen in werkprocessen, het testen van GPAI-capaciteiten en/of het werken aan AI-beleid.
108125

109-
{{< embed_pdf url="/pdf-files/events/activities/20260625_Masterclass_Benchmarking.pdf" width_mobile_pdf="12" width_desktop_pdf="6" >}}
126+
{{< embed_pdf url="/pdf-files/events/activities/20261103_Masterclass_Benchmarking.pdf" width_mobile_pdf="12" width_desktop_pdf="6" >}}
110127

111128
{{< dynamic_form_engine index="0" >}}
112129

3.23 MB
Loading
Binary file not shown.

0 commit comments

Comments
 (0)