Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 21 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,9 +190,29 @@ spreads, because a baseline is usually recorded for far longer than the run comp
`CompareWithExpected` tests one session against the 0.5 an unbiased generator gives. Both are two-tailed: a
one-tailed test would find a shift toward a target more easily, but only holds when the direction was
predicted before the readings were taken, which the application cannot know. Every path that cannot produce
a number — fewer than two readings, or readings with no spread — returns `Valid = false` rather than a
a number - fewer than two readings, or readings with no spread - returns `Valid = false` rather than a
probability, and `Significant` is false whenever the test did not run.

A verdict of nothing found means nothing at all unless a shift worth finding could have been seen, so the
analysis states what it could have seen. `SignificanceResult.DetectableDifference` is the half width of the
confidence interval for the pair, which is the test read the other way round: instead of asking whether the
shift that happened beats the noise, it asks how large a shift would have to be before it could. It comes
back with the test, worked out in `BuildResult` from the same critical value the probability comes from, so
the readings are walked once and the limit cannot disagree with the test it is quoted beside. Below
`SHIFT_OF_INTEREST` - one part in ten thousand, the order of the effect reported in the published work - the
sessions cannot speak to the question, and the verdict says so in the warning colour rather than reporting a
null result that reads as evidence of absence.

The three verdicts are different weights, which is what `VerdictWeights` carries. A shift found is notable
whatever the sensitivity, because it cleared the limit by being found at all. Nothing found from sessions
that could have found something is the ordinary outcome. Nothing found from sessions that could not is the
one that misleads, and is the only one that warns.

Note that the device and the simulator have different noise floors: a device reading averages 262,144 bits
and a simulated one 16,384, so the simulated spread is about four times wider and needs about sixteen times
the readings to pin its mean down as finely. That is why the limit is worked out from each session's own
spread rather than from a count of readings.

The verdict is stated in words under the table rather than left to be read off the numbers, and is
emphasised only when there is a shift to notice. `HistogramChart` plots each session as a percentage of its
own readings, not as counts: on counts the longer session stands taller in every bin and hides the shift the
Expand Down
4 changes: 2 additions & 2 deletions GeneratorForm.Designer.cs

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

86 changes: 74 additions & 12 deletions GeneratorForm.cs
Original file line number Diff line number Diff line change
Expand Up @@ -81,6 +81,18 @@ public enum RngGuiStates
RNG_GUI_STATES_SIZE // Keep at end
};

/// <summary>
/// How much attention a verdict is worth. Nothing found is the ordinary outcome of the experiment
/// and is not worth shouting about; a shift is worth noticing; and a session that could not have
/// shown the shift it was looking for is worth warning about, because its silence means nothing.
/// </summary>
public enum VerdictWeights
{
Ordinary = 0,
Notable = 1,
Warning = 2,
};

#endregion
#region Constructors

Expand Down Expand Up @@ -2450,32 +2462,62 @@ private void UpdateComparisonVerdict(List<double> baselineReadings, List<double>
bool bBothLoaded = ((null != baselineReadings) && (null != resultReadings));
if (false == bBothLoaded)
{
SetVerdict(m_sVERDICT_NEEDS_BOTH, false);
SetVerdict(m_sVERDICT_NEEDS_BOTH, VerdictWeights.Ordinary);
return;
}

SignificanceResult test = SignificanceTest.CompareMeans(baselineReadings, resultReadings);
if (false == test.Valid)
{
SetVerdict(m_sVERDICT_NOT_ENOUGH_DATA, false);
SetVerdict(m_sVERDICT_NOT_ENOUGH_DATA, VerdictWeights.Warning);
return;
}

double fShift = (Mean(resultReadings) - Mean(baselineReadings));
string sDirection = (0 <= fShift) ? m_sDIRECTION_HIGHER : m_sDIRECTION_LOWER;
string sProbability = FormatProbability(test);
string sFreedom = test.DegreesOfFreedom.ToString(m_sMOMENT_FORMAT);
string sShift = Math.Abs(fShift).ToString(m_sVALUE_FORMAT);

// How finely the two sessions between them pin a difference down. A verdict of nothing found
// means nothing at all unless a shift worth finding could have been seen, so the limit is
// stated either way rather than left for the reader to work out from the reading counts. It
// comes back with the test rather than being asked for separately, so the readings are walked
// once and the limit cannot disagree with the test it is quoted beside.
double fDetectable = test.DetectableDifference;
bool bSensitive = SignificanceTest.IsSensitiveEnough(fDetectable);
Comment thread
Copilot marked this conversation as resolved.
string sDetectable = fDetectable.ToString(m_sVALUE_FORMAT);
string sOfInterest = SignificanceTest.SHIFT_OF_INTEREST.ToString(m_sVALUE_FORMAT);

if (true == test.Significant)
{
SetVerdict($"Significant shift: the result sits {sDirection} than the baseline by " +
$"{Math.Abs(fShift).ToString(m_sVALUE_FORMAT)}. A shift this large would arise by " +
$"chance {sProbability} of the time (Welch's t, two-tailed, {test.DegreesOfFreedom.ToString(m_sMOMENT_FORMAT)} df).", true);
// The shift cleared the limit by having been found at all, so this reports the limit rather
// than warning about it, however short the sessions were
SetVerdict($"Significant shift: the result sits {sDirection} than the baseline by {sShift}. " +
$"A shift this large would arise by chance {sProbability} of the time " +
$"(Welch's t, two-tailed, {sFreedom} df). These sessions can show a difference of " +
$"{sDetectable} or larger.", VerdictWeights.Notable);
}
else if (true == bSensitive)
{
SetVerdict($"No significant shift. The result sits {sDirection} than the baseline by {sShift}, " +
$"but a shift that large would arise by chance {sProbability} of the time " +
$"(Welch's t, two-tailed, {sFreedom} df). These sessions can show a difference of " +
$"{sDetectable} or larger, so a shift of {sOfInterest} would have been found.",
VerdictWeights.Ordinary);
}
else
{
SetVerdict($"No significant shift. The result sits {sDirection} than the baseline by " +
$"{Math.Abs(fShift).ToString(m_sVALUE_FORMAT)}, but a shift that large would arise by " +
$"chance {sProbability} of the time (Welch's t, two-tailed, {test.DegreesOfFreedom.ToString(m_sMOMENT_FORMAT)} df).", false);
// Nothing was found and nothing could have been. This is the reading that misleads if it is
// left to stand on its own, so it is the one that carries the warning. The measurement is
// still reported in full first: the warning is about what the numbers can be taken to mean,
// not a reason to stop showing them.
SetVerdict($"No significant shift. The result sits {sDirection} than the baseline by {sShift}, " +
$"but a shift that large would arise by chance {sProbability} of the time " +
$"(Welch's t, two-tailed, {sFreedom} df). These sessions are too short to conclude " +
$"anything from that: they can only show a difference of {sDetectable} or larger, so " +
$"a shift of {sOfInterest}, the size this looks for, could be real here and still " +
$"never reach significance. Record for longer.", VerdictWeights.Warning);
}
}

Expand All @@ -2484,12 +2526,32 @@ private void UpdateComparisonVerdict(List<double> baselineReadings, List<double>
/// ordinary outcome and is not worth shouting about.
/// </summary>
/// <param name="sVerdict">IN - The wording to show</param>
/// <param name="bSignificant">IN - Whether the verdict reports a shift worth noticing</param>
private void SetVerdict(string sVerdict, bool bSignificant)
/// <param name="weight">IN - How much attention the verdict is worth</param>
private void SetVerdict(string sVerdict, VerdictWeights weight)
{
m_VerdictLabel.Text = sVerdict;
m_VerdictLabel.Font = bSignificant ? m_VerdictFontBold : m_VerdictFontRegular;
m_VerdictLabel.ForeColor = bSignificant ? UiPalette.CardText : UiPalette.MutedText;

// Anything other than the ordinary outcome is set in the heavier face, so a verdict that needs
// reading is not the same weight as the one that says nothing happened
bool bOrdinary = (VerdictWeights.Ordinary == weight);
m_VerdictLabel.Font = bOrdinary ? m_VerdictFontRegular : m_VerdictFontBold;

// A warning takes the severity colour, which stands aside under a high contrast scheme and
// leaves the wording to carry it
switch (weight)
{
case VerdictWeights.Warning:
m_VerdictLabel.ForeColor = StatusPalette.WarningText;
break;

case VerdictWeights.Notable:
m_VerdictLabel.ForeColor = UiPalette.CardText;
break;

default:
m_VerdictLabel.ForeColor = UiPalette.MutedText;
break;
}
}

/// <summary>
Expand Down
13 changes: 13 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,19 @@ directly: mean, standard deviation, skewness and kurtosis for each, the differen
a histogram of the two distributions overlaid. Statistics are calculated with
[Math.NET Numerics](https://numerics.mathdotnet.com/).

### Enough readings to answer

A session that found nothing has only said something if it could have found something. The verdict therefore
states the smallest difference the two sessions could show as significant, and warns when that is coarser
than the shift being looked for — one part in ten thousand, a generator running at 0.5001 rather than 0.5,
which is the order of the effect reported in the published work.

A minute or so of recording from the device is enough to speak to a shift that size. Below that, a real
shift can sit in the readings and never reach significance, and a verdict of *no significant shift* would
be reporting the length of the session rather than anything about the generator. The limit is worked out
from each session's own spread rather than from a count of readings, because a simulated reading averages
far fewer bits than a device one and is correspondingly noisier.

### Targets

A session can record a target value of `0` or `1` — the outcome the operator is attempting to influence —
Expand Down
Loading