Skip to content

[flaky-ci] Stabilize BenchmarkDotNet device test results - #12718

Merged
simonrozsival merged 5 commits into
mainfrom
simonrozsival-benchmarkdotnet-device-stability
Sep 10, 2026
Merged

simonrozsival merged 5 commits into
mainfrom
simonrozsival-benchmarkdotnet-device-stability

Conversation

@simonrozsival

Copy link
Copy Markdown
Member

Fixes #12704

Tracker: #12704

Focused tracker: N/A

Why

InstallAndRunTests.DotNetRunBenchmarkDotNet failed in builds 1573240 and 1580340 even though the app built, installed, launched, ran BenchmarkDotNet, returned reports=1, and finished instrumentation successfully. Only app logcat events were missing. Local repeated validation reproduced a run with exit code 0 and complete successful instrumentation results while both app logcat markers were absent, confirming that logcat delivery—not build, installation, launch, benchmark startup, device performance, parsing, or timeout behavior—was nondeterministic.

Microsoft.Android.Run stops its logcat reader when am instrument exits, so assertions must not depend on app logcat arriving before that shutdown.

What changed

  • Return the exact forwarded args and greeting values in the instrumentation result bundle.
  • Return benchmark discovery, report, successful-report, and critical-validation-error counts.
  • Assert the structured instrumentation results as the authoritative integration outcome.
  • Keep both app logcat markers as diagnostics only, with a concise milestone summary in test output.
  • Preserve the real BenchmarkDotNet 0.15.8 in-process run, instrumentation-code and process-exit checks, and the existing timeout.

Validation

  • ./dotnet-local.sh build tests/MSBuildDeviceIntegration/MSBuildDeviceIntegration.csproj -c Debug -v:minimal

  • DotNetRunBenchmarkDotNet passed three consecutive emulator runs.

  • A pre-fix repetition reproduced missing app logcat with successful structured benchmark results.

  • Useful description of why the change is necessary

  • Links to issues fixed

  • Focused device integration coverage

simonrozsival and others added 2 commits September 8, 2026 11:58
Use instrumentation results for benchmark and argument assertions so dropped logcat output cannot fail the test. Add structured diagnostics for discovery, report success, and validation errors.\n\nCo-authored-by: Copilot App <[email protected]>
@simonrozsival simonrozsival added the flaky-ci Intermittent CI failures and work to improve CI reliability label Sep 8, 2026
@simonrozsival

Copy link
Copy Markdown
Member Author

/review

@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

✅ Android PR Reviewer completed successfully!

Generated by Android PR Reviewer for #12718

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Needs Changes

The structured instrumentation-result approach correctly removes the test dependency on nondeterministic logcat delivery, and the added assertions cover argument forwarding, benchmark discovery, report success, and validation errors. I found 0 errors, 0 warnings, and 1 suggestion (inline) to improve parser failure diagnostics.

CI is currently red: Azure build 1586987 reports failures in Package Tests macOS > Tests > APKs 1 and MAUI Tests MAUI Integration. The remaining 41 checks passed. I could not determine whether those two failures are flaky/infrastructure-related because Azure CLI could not initialize in this runner, so CI must be investigated or rerun before merge.

Generated by Android PR Reviewer for #12718 · gpt56 · 260.3 AIC · ⌖ 3.76 AIC · ⊞ 27.7K
Comment /review to run again

Comment thread tests/MSBuildDeviceIntegration/Tests/InstallAndRunTests.cs Outdated
Report missing keys separately from malformed integer values so output-path failures remain actionable.

Co-authored-by: Copilot App <[email protected]>
@simonrozsival

Copy link
Copy Markdown
Member Author

/review

@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

✅ Android PR Reviewer completed successfully!

Generated by Android PR Reviewer for #12718

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Needs Changes

Findings: 0 errors · 1 warning · 0 suggestions.

The structured instrumentation bundle is a sound replacement for timing-sensitive logcat assertions, and the added success/report/validation checks strengthen the test. One diagnostic-ordering issue remains: exception runs can fail while parsing absent success keys before the test checks the instrumentation result code.

CI build #1587831 is currently red in the Linux, macOS, and Windows build lanes. The Azure CLI could not read the public timeline in this runner because its configured profile location is not writable, so I could not determine whether those failures are related to this PR.

Generated by Android PR Reviewer for #12718 · gpt56 · 222.7 AIC · ⌖ 13.6 AIC · ⊞ 25.7K
Comment /review to run again

Comment thread tests/MSBuildDeviceIntegration/Tests/InstallAndRunTests.cs
simonrozsival and others added 2 commits September 9, 2026 11:09
Check the instrumentation and process completion status before parsing success-only result values, and surface the instrumentation error bundle when execution is canceled.

Co-authored-by: Copilot App <[email protected]>
@simonrozsival
simonrozsival marked this pull request as ready for review September 10, 2026 10:44
Copilot AI lite review requested due to automatic review settings September 10, 2026 10:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new instrumentation-result parsing trims extracted values, which can mutate the “authoritative” forwarded args/greeting and undermine the stated goal of returning exact values.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review tier: Lite
Findings: 1 Medium severity

New issues introduced by this change (1)
Severity Finding
Medium severity tests/​MSBuildDeviceIntegration/​Tests/​InstallAndRunTests.cs — TryParseInstrumentationStringResult() trims the whole line and the extracted value. That can…
What changed in this PR

This PR updates the DotNetRunBenchmarkDotNet MSBuild device integration test to avoid flakiness caused by nondeterministic delivery of app logcat markers when am instrument exits, by treating structured instrumentation bundle results as the authoritative signal.

Changes:

  • Extend the in-app Instrumentation to emit benchmark/run metadata (args/greeting, benchmark/report counts, validation error counts) via INSTRUMENTATION_RESULT.
  • Update the test to assert on INSTRUMENTATION_RESULT keys instead of relying on app logcat markers, while still printing logcat-marker presence as diagnostics.
  • Refactor parsing helpers to support both integer and string instrumentation result values.
File Description
tests/​MSBuildDeviceIntegration/​Tests/​InstallAndRunTests.cs Makes BenchmarkDotNet device-test assertions rely on instrumentation results rather than app logcat, and adds additional structured result keys for stability/diagnostics.

Comment on lines 3611 to 3615
foreach (var rawLine in output.Split ('\n')) {
var line = rawLine.Trim ();
if (line.StartsWith (prefix, StringComparison.Ordinal)) {
var valueStr = line.Substring (prefix.Length).Trim ();
if (int.TryParse (valueStr, out int value))
return value;
return line.Substring (prefix.Length).Trim ();
}
@simonrozsival
simonrozsival merged commit 4e78319 into main Sep 10, 2026
45 checks passed
@simonrozsival
simonrozsival deleted the simonrozsival-benchmarkdotnet-device-stability branch September 10, 2026 20:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

flaky-ci Intermittent CI failures and work to improve CI reliability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CI flakiness inventory: Aug 11-Sep 8, 2026

3 participants