## Summary of the Pull Request Adds a `PowerRename.UITests.Next` suite powered by winappcli and automates all 18 scenarios from #40663. The suite covers PowerRename settings, search and replace behavior, regular expressions, formatting and filtering options, file-list interactions, and both classic and Windows 11 context-menu workflows. The PR also stabilizes shared `UITestAutomation.Next` runner lifetimes and settings restoration, adds automation IDs for the original and renamed counters, and prepares unsigned CI builds for PowerRename shell testing. CI now signs the sparse context-menu MSIX and the runner/Settings IPC companions with a disposable machine-trusted test identity. ## PR Checklist - [x] Closes: #40663 - [x] **Communication:** I've discussed this with core contributors already. If the work hasn't been agreed, this work might be rejected - [x] **Tests:** Added/updated and all pass - [x] **Localization:** All end-user-facing strings can be localized - [x] **Dev docs:** Added/updated - [x] **New binaries:** Not applicable; no new shipped product binaries are added - [x] JSON for signing: Not applicable - [x] WXS for installer: Not applicable - [x] YML for CI pipeline: The new UI-test project is discovered through the existing `*UITest*.csproj` pipeline flow and is registered in `PowerToys.slnx` - [x] YML for signed pipeline: Not applicable - [x] **Documentation updated:** Not applicable; there are no user-facing behavior or documentation changes ## Detailed Description of the Pull Request / Additional comments ### PowerRename UI tests - Adds a 25-case `PowerRename.UITests.Next` executable covering all 18 checklist items from #40663. - Exercises classic context-menu registration on Windows 10 and Windows 11. - Exercises the signed Windows 11 tier-1 context menu, including icon visibility and real invocation with an Explorer selection. - Covers search/replace preview and application, text formatting, file/folder/subfolder inclusion, filename/extension scope, enumeration, case sensitivity, match-all behavior, regular expressions, file timestamps, Boost syntax, MRU autocomplete, persisted values, and file-list selection/filtering. - Preserves the existing legacy tests. ### Test reliability and automation hooks - Reuses one runner/Settings lifetime across the complete PowerRename suite to avoid repeated cold launches on constrained agents. - Retains and restores global settings for the full class lifetime and verifies that both Settings and the runner remain healthy. - Propagates scope and PowerRename cleanup failures instead of silently leaking process or profile state. - Adds `OriginalCount` and `RenamedCount` automation IDs. These are automation-only metadata and do not change the visible UI. - Uses stable preview samples, exact count targeting, authoritative Explorer selection, readable classic-menu inventories, and live UIA visibility for popup items. ### Unsigned CI build support - Extends the existing sparse-package test signer with required Authenticode companion files. - PowerRename jobs sign `PowerToys.exe` and `PowerToys.Settings.exe` with the same disposable machine-trusted test identity used for sparse MSIX packages. - This preserves Release IPC authentication while allowing Settings module-toggle commands to work on unsigned PR builds. - Windows 11 and ARM64 PowerRename jobs require a validly signed `PowerRenameContextMenuPackage.msix` before tests start. ## Validation Steps Performed ### Builds and discovery - `PowerRename.UITests.Next` Debug x64 build: passed - `PowerRename.UITests.Next` Debug ARM64 cross-build: passed - `PowerRenameUI` Release x64 build: passed - `UITestAutomation.Next.UnitTests` Debug x64 build: passed - `UITestAutomation.Next.UnitTests`: 16/16 passed - Microsoft.Testing.Platform discovery: 25 unique PowerRename test cases ### Local VM matrix All runs used the complete unfiltered `TestCategory=PowerRename` suite and restored the standard user's settings file byte-for-byte. | Guest | Profile | Result | |---|---|---:| | Windows 10 x64 | Default, 4 vCPU / 8 GB | 25/25 | | Windows 10 x64 | Constrained, 1 vCPU / 4 GB | 25/25 | | Windows 11 x64 | Default, 4 vCPU / 8 GB | 25/25 | | Windows 11 x64 | Constrained, 1 vCPU / 4 GB | 25/25 | ### Azure DevOps UI Test Automation - Final build: [155646235 / 20260824.1](https://dev.azure.com/microsoft/Dart/_build/results?buildId=155646235) - Source revision: `4496683104aea622b30179909bd6e95a17d5500f` - ARM64: 25/25 passed - Windows 10 x64: 25/25 passed - Windows 11 x64: 25/25 passed - Total: 75/75 passed, with zero failed, skipped, not-executed, or unanalyzed results - ARM64 and x64 Release product builds succeeded and published their normal artifacts - Signing steps verified the PowerRename sparse MSIX where applicable and both IPC companion executables on every PowerRename test job <img width="372" height="314" alt="image" src="https://github.com/user-attachments/assets/bc303673-0325-4f85-8d70-ef93110baf5b" />
20 KiB
Azure DevOps UI-test CI agentic loop
This is an internal post-local-validation workflow. It assumes the selected UITest project already passed the complete local matrix required by ui-tests-local-vm, including default and constrained profiles.
0. Use the existing Azure CLI session
All Azure DevOps operations in this workflow use PowerShell 7, the existing Azure CLI sign-in, and Azure DevOps REST APIs. This avoids per-call authentication prompts and works for definitions, builds, timelines, logs, preview/queue, stage retry/cancel, test results, artifacts, and result attachments.
Required one-command readiness gate
Before the first Azure operation in each agent session, run the preflight in a fresh PowerShell 7 process:
pwsh -NoLogo -NoProfile -File `
.github\skills\ui-tests-pipeline-ci\scripts\Test-AzureDevOpsSetup.ps1
Do not proceed unless it exits 0, reports Ready: true, and every required check is PASS. The
default probe uses refs/heads/main, module FancyZones.UITests.Next, and dynamically selects one
of the ten newest completed pipeline builds for build/log/artifact checks plus the first test-bearing
build in that set for Azure Test checks. The JSON reports these separately as ProbeBuildId and
ProbeTestBuildId. It creates no build and changes no Azure or repository state.
Use explicit probe inputs when diagnosing a particular branch or known build:
pwsh -NoLogo -NoProfile -File `
.github\skills\ui-tests-pipeline-ci\scripts\Test-AzureDevOpsSetup.ps1 `
-ProbeBranch refs/heads/<branch> `
-ProbeModule <Module.UITests> `
-ProbeBuildId <KNOWN_COMPLETED_BUILD_ID>
For the check inventory, precise capability claims, and first-time remediation, read setup-preflight.md only when setup fails or the user asks about readiness. The preflight deliberately performs no mutation.
The agent never starts an interactive sign-in or installs tools. If preflight fails, stop and report
the exact failed check. The user performs any required setup outside the agent, then the agent reruns
the same preflight. Never pass credentials through chat or run az login from the agent.
Use the REST helper after preflight
Dot-source the bundled helper for actual work only after preflight passes:
. .\.github\skills\ui-tests-pipeline-ci\scripts\AzureDevOps.ps1
The helper obtains a token for Azure DevOps resource
499b84ac-1321-427f-aa17-267ca6975798 on each request. Azure CLI serves it from its cache and
silently refreshes it when needed. The token and authorization header stay in memory and are cleared
after every request. Repeated reads and actual pipeline queueing were verified without prompts on
2026-08-20.
Never run az login through an agent, request credentials, print or persist a token/header, enable
command tracing around authentication, or commit downloaded internal evidence. If
the preflight fails, ask the user to resolve its exact failed check and stop with an access blocker.
Do not fall back to another transport.
Invoke-AzDevOpsRest accepts a project-relative REST path and returns { Body, Headers }:
$build = (Invoke-AzDevOpsRest -Uri '_apis/build/builds/123?api-version=7.1').Body
Get-AzDevOpsPagedValues follows every x-ms-continuationtoken response header. Use it whenever an
endpoint can paginate; never infer completeness from one page.
REST writes mutate Azure state. Queue, cancel, retry, or approve only when the user's CI request authorizes that action. The absence of a per-call authentication dialog is not authorization.
1. Prove the local gate
Before touching Azure DevOps, record:
- Exact pushed branch and commit SHA.
- Clean x64 and ARM64 builds where applicable.
- Complete suite on Windows 10 and Windows 11 under the default profile.
- Complete suite under
Constrained(1 vCPU, 4 GB). - Windows 11 ARM64 guest evidence on a Windows-on-ARM host when applicable.
- Zero skipped, inconclusive, or not-executed tests and zero export errors.
Do not substitute a focused run for full sign-off. If a required local environment is unavailable, stop and ask the user before consuming CI.
Derive uiTestModules from the exact .csproj filename without .csproj, for example:
FancyZonesEditor.UITests.Next
The list must be non-empty and contain only the projects currently being changed.
2. Discover the pipeline and serialize the branch
Discover rather than assuming IDs:
$branch = 'refs/heads/<current-branch>'
$definitionName = [Uri]::EscapeDataString('UI Test Automation')
$definitions = (Invoke-AzDevOpsRest -Uri `
"_apis/build/definitions?name=$definitionName&api-version=7.1").Body.value
$enabled = @($definitions | Where-Object queueStatus -EQ 'enabled')
if ($enabled.Count -ne 1) {
throw "Expected one enabled UI Test Automation definition; found $($enabled.Count)."
}
$pipelineId = [int]$enabled[0].id
$encodedBranch = [Uri]::EscapeDataString($branch)
$history = Get-AzDevOpsPagedValues -Uri `
"_apis/build/builds?definitions=$pipelineId&branchName=$encodedBranch&queryOrder=queueTimeDescending&%24top=100&api-version=7.1"
$active = @($history.Items | Where-Object status -IN @('notStarted', 'inProgress', 'postponed', 'cancelling'))
Require every returned build's sourceBranch to equal the exact target ref. Normally only one run
for that branch may be active. Other branches do not block queueing and must not be canceled.
Adopt an active run only when its source SHA, selected platforms, modules, buildSource, and reused
build ID exactly match the checkpoint. A mismatching same-branch run is a blocker. Wait, or cancel
only a superseded run that the user owns or explicitly asked to stop:
Invoke-AzDevOpsRest `
-Uri "_apis/build/builds/${buildId}?api-version=7.1" `
-Method Patch `
-Body @{ status = 'cancelling' }
Confirm it becomes terminal completed/canceled; cancelling still occupies the branch slot.
A narrow exception allows parallel same-branch execution only when the user explicitly authorizes a supplemental run for an architecture whose product build failed while another architecture remains active. Use the same pushed SHA and modules, select only the failed architecture, and checkpoint both build IDs. The supplemental run remains part of the attempt ledger and never erases the original failure.
3. Choose buildNow or specificBuildId
| Situation | buildSource |
specificBuildId |
|---|---|---|
| First run in the sequence | buildNow |
xxxx |
| Product/runtime/common/pipeline files changed | buildNow |
xxxx |
| Prior product build failed, was incomplete, or lacks one selected-platform artifact | buildNow |
xxxx |
| Only selected UITest project files changed and every selected product artifact succeeded | specificBuildId |
Prior numeric build ID |
For reuse, prove all of the following:
- Previous target-branch run is terminal.
- Product build stages succeeded for every selected platform.
- Artifact names include normal
build-x64-Releaseand/orbuild-arm64-Releaseartifacts as selected.build-<platform>-Release-failure-<attempt>is diagnostic and never reusable. git diff --name-only <previous-sourceVersion>..HEADis confined to selected UITest project directories. Any shared, product, pipeline, dependency, or installer change requiresbuildNow.- The revision is pushed.
Use numeric build IDs such as 154921069, never display numbers such as 20260814.4.
4. Preview and queue
Default parameters:
| Parameter | Value |
|---|---|
buildPlatforms |
- arm64\n- x64 |
enableMsBuildCaching |
false |
useVSPreview |
false |
useLatestWebView2 |
false |
buildSource |
Decision from section 3 |
specificBuildId |
xxxx or prior numeric build ID as a string |
uiTestModules |
Bracketed non-empty list, e.g. [KeyboardManager.UITests] |
Do not silently remove a platform. x64 expands to Windows 10 and Windows 11 jobs; arm64 expands
to the ARM64 job.
Preview first with the exact branch and string-valued template parameters:
$templateParameters = @{
buildPlatforms = "- arm64`n- x64"
enableMsBuildCaching = 'false'
useVSPreview = 'false'
useLatestWebView2 = 'false'
buildSource = 'buildNow'
specificBuildId = 'xxxx'
uiTestModules = '[KeyboardManager.UITests]'
}
$request = @{
previewRun = $true
resources = @{ repositories = @{ self = @{ refName = $branch } } }
templateParameters = $templateParameters
}
$preview = (Invoke-AzDevOpsRest `
-Uri "_apis/pipelines/${pipelineId}/runs?api-version=7.1-preview.1" `
-Method Post `
-Body $request).Body
Inspect finalYaml. Require only the expected Build_<platform> and test stages, exact module
assignments, no unrequested platform, and no prior build ID for buildNow. Preview creates no real
build and does not consume an attempt.
Immediately repeat the full branch preflight from section 2. If clear, queue by changing only:
$request.previewRun = $false
$run = (Invoke-AzDevOpsRest `
-Uri "_apis/pipelines/${pipelineId}/runs?api-version=7.1-preview.1" `
-Method Post `
-Body $request).Body
Record id, name, web link, branch, resolved repository version, and echoed parameters. Query the
returned build ID and require sourceVersion to equal the pushed SHA. If it differs, cancel before
tests and stop.
Queue and active-run discovery are not atomic. Immediately list every active exact-branch run again, following all continuation tokens. The oldest matching run owns the normal branch slot. Cancel a younger run created by this agent when safe; never cancel someone else's run. Every actual queue, including a reconciled duplicate, counts in the ledger.
Persist the checkpoint
Store in session/task state, never the repository:
Pipeline: UI Test Automation / <definition ID>
Build: <build ID> / <build number> / <web link>
Branch: <exact refs/heads/... ref>
Source SHA: <40-character commit>
Attempt: <n>/3
Build source: <buildNow|specificBuildId> [reused build ID]
Platforms: <selected platforms>
Modules: [<exact project stems>]
State: <notStarted|inProgress|...>
After any interruption, scheduled wake, or user turn, query every exact checkpointed build ID before other Azure work. Never rediscover by adopting the newest run.
One-hour scheduled continuation
Do not leave a nonterminal build with a passive handoff. Full builds normally take 60-90 minutes and test-only validation about 40 minutes. Arm one one-shot host wake-up at a time; after each wake, query exact build IDs and re-arm for one hour only if work remains nonterminal.
Monitoring state is scoped to exact build IDs. If the user reports a tracked build's status and asks to stop its now-obsolete watcher, remove only that build's task/marker. Do not treat that as a standing opt-out: every later build queued by the agent must immediately receive a fresh one-shot continuation and remain monitored until terminal unless the user explicitly opts out for that new build.
The scheduled action only wakes the agent. Do not put az, a token, build parameters, or any Azure
request in it. On Windows without a native agent scheduler:
- Persist build IDs, unique task name, marker under
$env:TEMP, scheduled UTC/local time, and watcher terminal ID in session state. - Register a current-user, limited one-shot task for
(Get-Date).AddHours(1)whose only action writes the marker. Use an encoded PowerShell command for deterministic quoting. - Start one asynchronous
FileSystemWatcherterminal for the marker. Do not useStart-Sleep, a timer loop, or repeated Azure requests. - On notification, remove marker and task, then query exact IDs through
Invoke-AzDevOpsRest. - Re-arm only after an authenticated status query confirms a build remains nonterminal.
5. Monitor exact builds
Use these endpoints through Invoke-AzDevOpsRest:
| Evidence | REST path |
|---|---|
| Build status/source | _apis/build/builds/<BUILD_ID>?api-version=7.1 |
| Stage/job timeline and issues | _apis/build/builds/<BUILD_ID>/timeline?api-version=7.1 |
| Log metadata | _apis/build/builds/<BUILD_ID>/logs?api-version=7.1 |
| Narrow log content | _apis/build/builds/<BUILD_ID>/logs/<LOG_ID>?startLine=<N>&endLine=<N>&api-version=7.1 |
| Pipeline artifacts | _apis/build/builds/<BUILD_ID>/artifacts?api-version=7.1 |
| Test runs for build | _apis/test/runs?buildUri=vstfs%3A%2F%2F%2FBuild%2FBuild%2F<BUILD_ID>&api-version=7.1 |
| Paged test results | _apis/test/Runs/<RUN_ID>/results?%24top=1000&%24skip=<N>&api-version=7.1 |
Read build status, then timeline, newest log timestamps, relevant narrow logs, every non-passing test
result, and artifacts. Page test results by $skip until fewer than 1,000 remain. Classify every
outcome other than Passed, including failed, aborted, error, timeout, not-executed, inconclusive,
blocked, warning, not-applicable, and paused.
Do not busy-poll. Fresh timeline/log timestamps are progress.
| REST status/result | Action |
|---|---|
notStarted, inProgress, postponed, cancelling |
Continue monitoring exact ID |
completed + succeeded |
Verify selected tests executed and every result passed |
completed + partiallySucceeded |
Treat as failure until warnings/results are understood |
completed + failed |
Diagnose logs, tests, screenshots, and recordings |
completed + canceled |
Record who/why; not a product/test failure |
completed + abandoned |
Infrastructure/administrative termination |
Retry or cancel one stage
Use the YAML stage key such as Build_x64, not its display name. Retry only a terminal failed stage:
$stageRefName = 'Build_x64'
Invoke-AzDevOpsRest `
-Uri "_apis/build/builds/${buildId}/stages/${stageRefName}?api-version=7.1-preview.1" `
-Method Patch `
-Body @{ state = 'retry'; forceRetryAllJobs = $false }
REST success is not proof that retry started. Re-read the timeline until the stage/job attempt increments and becomes pending/in-progress. Retry materialization can be delayed. Do not queue a supplemental run until that check establishes the update was rejected or remained a no-op.
Use state = 'cancel' to cancel one known stage. Use the build PATCH from section 2 for the whole
run. Re-read status after every mutation.
6. Diagnose failures and collect evidence
First determine whether selected test stages ran. An empty failed-result query can also mean a product build failed and dependent tests were skipped.
Product build failed before tests
For each selected platform:
- Read stage/job timeline issues as routing data.
- Locate
Build Release_<platform>and read a narrow log tail around the first##[error], compiler error, MSBuild error summary, test failure, or nonzero exit. Report the first actionable cause, not the later generic task-exit message. - Compare artifacts:
build-<platform>-Releaseis the normal product artifact.build-<platform>-Release-failure-<attempt>is diagnostics only.
- Confirm dependent test stages are skipped.
- State that no screenshot/video exists when no test ran.
- Never reuse a partial product build with
specificBuildId.
Download a useful failure artifact to a temporary directory with the already-installed extension:
az pipelines runs artifact download `
--organization https://dev.azure.com/microsoft `
--project Dart `
--run-id <BUILD_ID> `
--artifact-name <FAILURE_ARTIFACT_NAME> `
--path <TEMP_DIRECTORY> `
--only-show-errors
Do not install tools or commit evidence. Inspect text logs first; open .binlog only with an already
available Structured Log Viewer.
UI test stages ran
- Query all non-passing results and preserve
(runId, resultId)because titles repeat by platform. - Resolve platform from the owning job/log, never result order.
- Read error, stack, duration, assembly, and
Standard_Console_Output.log. - Open failure screenshot and
recording_*.mp4before changing code. - Apply the authoritative-signal analysis from the local VM loop.
Construct an attachments-pane link for every failed result:
https://microsoft.visualstudio.com/Dart/_build/results?buildId=<BUILD_ID>&view=ms.vss-test-web.build-test-results-tab&runId=<RUN_ID>&resultId=<RESULT_ID>&paneView=attachments
Download result attachments without browser help
Azure Test result attachments are separate from pipeline artifacts. List and download them directly:
$buildId = <BUILD_ID>
$runId = <RUN_ID>
$resultId = <RESULT_ID>
$destination = Join-Path $env:TEMP "PowerToys-CI-$buildId-$runId-$resultId"
New-Item -ItemType Directory -Path $destination -Force | Out-Null
$resultBase = "_apis/test/Runs/${runId}/Results/${resultId}"
$attachments = (Invoke-AzDevOpsRest `
-Uri "$resultBase/attachments?api-version=7.1-preview.1").Body.value
$selected = @($attachments | Where-Object {
$_.fileName -eq 'Standard_Console_Output.log' -or
$_.fileName -like 'failure-*.png' -or
$_.fileName -like 'recording_*.mp4'
})
if (-not ($selected.fileName -contains 'Standard_Console_Output.log')) {
throw 'Standard_Console_Output.log is missing from the failed result.'
}
foreach ($attachment in $selected) {
$target = Join-Path $destination ([IO.Path]::GetFileName([string]$attachment.fileName))
Invoke-AzDevOpsRest `
-Uri "$resultBase/Attachments/$($attachment.id)?api-version=7.1-preview.1" `
-OutFile $target | Out-Null
[pscustomobject]@{
AttachmentId = $attachment.id
FileName = Split-Path $target -Leaf
Bytes = (Get-Item $target).Length
Sha256 = (Get-FileHash $target -Algorithm SHA256).Hash
Path = $target
}
}
Read the console log and relevant product logs directly, inspect visual evidence, and record bytes plus SHA-256. Never ask the user to fetch files that this path can download.
Pipeline artifact metadata contains resource.downloadUrl; include useful authenticated links in
reports. Do not confuse multi-gigabyte product artifacts with Azure Test recordings.
7. Iterate, maximum three runs
Maintain this ledger:
| Attempt | Build ID / number | Source SHA | Build source | Product build | Tests | Failure signature | Evidence links | Progress |
|---|---|---|---|---|---|---|---|---|
| 1/3 | buildNow |
Baseline | ||||||
| 2/3 | ||||||||
| 3/3 |
Every actual queued build counts, including infrastructure failures and canceled duplicates. Preview runs do not. A supplemental architecture run also counts unless the user explicitly authorizes an exception after seeing the ledger.
Before another attempt:
- State one falsifiable hypothesis from logs and visual evidence.
- Make the smallest relevant fix without weakening assertions.
- Build and rerun the focused test locally.
- Rerun affected full default and constrained suites.
- Commit and push.
- Confirm current run terminal or explicitly handle the supplemental-architecture exception.
- Re-evaluate section 3.
Progress means fewer failures/platforms, the original test passing, a later authoritative state, or a broad timeout narrowed to actionable evidence. Error-text churn is not progress. After run 3 or three no-progress runs, stop and ask the user unless they explicitly authorize an exception.
8. Report the result
Include:
- Build link, numeric ID, display number, pipeline ID, branch, and exact SHA.
- Attempt number,
buildSource, reused ID, platforms, and modules. - Current/terminal status and per-platform counts.
- Every non-passing test and first actionable root-cause line.
- One attachments link per failed result, or an explicit no-test/no-video statement.
- Useful pipeline artifact links without confusing them with recordings.
- Whether next action is wait, local fix, retry, success, or escalation.