Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
73 changes: 73 additions & 0 deletions .claude/skills/watch-gradle-tests/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
---
name: watch-gradle-tests
description: Watch a Gradle instrumented-test run to a verdict without hand-writing greps. Use whenever running :app:testEmulatorDebugAndroidTest or :app:connectedDebugAndroidTest in the background.
---

# Watching a test run to a verdict

The emulator suite takes upwards of ten minutes, so it gets backgrounded, so something has to
report on it. **Do not write that something fresh each time.** Use the script.

```bash
# 1. Start the run, redirected to a file.
./gradlew :app:testEmulatorDebugAndroidTest --console=plain > tmp/run.log 2>&1

# 2. Prove the watcher sees the log before trusting it. One command, always.
python .claude/skills/watch-gradle-tests/watch_tests.py tmp/run.log --once

# 3. Arm a Monitor on the same script with no --once.
python .claude/skills/watch-gradle-tests/watch_tests.py tmp/run.log
```

Each line it prints is one event: `FAILED <Class>.<method>` as each failure appears,
`STALLED …` if the log stops growing, and `FINISHED` + `VERDICT` at the end.

## Why this exists

Three hand-written monitors in one afternoon each matched **nothing**, and each looked like a
green run:

| What was written | Why it matched nothing |
| --- | --- |
| `grep "(0 skipped)"` | the run had 5 skipped |
| `grep "\S+ \[testEmulator\]"` | there is no space before `[testEmulator]` |
| `... \| head -25` | truncated before the verdict, and `head`'s exit code hid it |

All three exited 0. **Silence from a monitor is indistinguishable from silence from a healthy
run**, so a red suite was reported as green until the log was read by hand. That is the whole
argument for a script: the patterns get fixed once, and `--once` proves they still match before
anything depends on them.

## What it knows that a grep does not

- **The XML has the last word.** The console can truncate, interleave, or count a retry
(`605/600` is a real line from this repo). `verdict_from_xml` reads
`app/build/outputs/androidTest-results/`, and **scopes to the newest run's own directory** —
AGP keeps `connected/` and `managedDevice/` side by side and clears neither, so summing
everything reported a 22-test run as 74. It always prints the results' age, because a verdict
that is quietly ten minutes old is the same bug wearing a coat.
- **No results at all is not zero failures.** A run that dies before any test reports writes no
XML, and reading that as success is exactly how a red build gets called green.
- **A stall is a result.** No output for eight minutes gets reported, with a diagnosis, rather
than looking like a slow test for another twenty.

## The two silent hangs, which need opposite fixes

The script separates them by whether *any* test has reported. Do not guess between them — the
distinction is in the output.

| Symptom | Cause | Fix |
| --- | --- | --- |
| No test ever reports, no managed AVD in `adb devices` | leaked managed-device slots — `MDLockCount` above zero with nothing running | `./gradlew --stop && rm ~/.android/avd/gradle-managed/active_gradle_devices` |
| Tests ran, then stopped mid-suite | the device died under them | `adb logcat -b crash` — the emulator's Bluetooth stack aborting has done this here |
| `connectedDebugAndroidTest` fails to install or hangs at once | stale ADB bridge inside the Gradle daemon | restart the emulator **and** `./gradlew --stop` — either alone leaves it broken |

Killing a run is what leaks a slot, and every `Ctrl-C` adds one. Prefer letting a run finish; if
you must kill one, clear the count in the same breath rather than meeting it next time.

## Related

- `AGENTS.md`, "Building and testing" — the managed device, and why it beats a hand-started
emulator.
- `.claude/skills/watch-pr/` — the same discipline for CI: prove the poll body emits before
arming a Monitor on it.
243 changes: 243 additions & 0 deletions .claude/skills/watch-gradle-tests/watch_tests.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,243 @@
#!/usr/bin/env python
"""
Watch a Gradle instrumented-test run and say what happened.

Written because hand-rolled greps kept reporting nothing and being believed. Three separate
monitors in one afternoon matched no lines each - one required ``(0 skipped)`` against a run with
five skipped, one required a space before ``[testEmulator]`` that is not there, one truncated
before the verdict with ``head``. Every one of them exited quietly, and quiet reads as green.

So the patterns live here, once, with ``--once`` to prove they match before anything relies on
them. Nothing about a run should be matched by a regex typed fresh at the call site.

Usage::

python watch_tests.py <logfile> --once # print what the log says now, and exit
python watch_tests.py <logfile> # poll, one line per new event, exit when done

Every line printed is an event worth a notification. Silence means nothing new, and only after
``--once`` has shown the patterns matching something.
"""

from __future__ import annotations

import argparse
import glob
import os
import re
import sys
import time
import xml.etree.ElementTree as ET

# **Anchored on the word FAILED, not on the decorations around it.** The name is followed
# immediately by `[deviceName]` with no space, and the line carries ANSI colour codes; both have
# already broken a hand-written pattern.
FAILED_LINE = re.compile(r"^(dev\.wander\S*) > (\S+?)\[")

# `(N skipped)` is not always `(0 skipped)`. Capture the numbers rather than matching a shape.
PROGRESS = re.compile(r"Tests (\d+)/(\d+) completed\. \((\d+) skipped\) \((\d+) failed\)")

# **MULTILINE, or `^` only matches the start of the whole file** and this reports "not yet"
# forever on a finished run. The same omission has already shipped once in redact.py.
TERMINAL = re.compile(r"^BUILD (SUCCESSFUL|FAILED)|^FAILURE: ", re.MULTILINE)

# Where AGP leaves the authoritative answer. The console can be truncated, retried or interleaved;
# these cannot.
RESULT_XML = "app/build/outputs/androidTest-results/**/*.xml"

LOCK_FILE = os.path.expanduser("~/.android/avd/gradle-managed/active_gradle_devices")


def failures_in(log: str) -> list[str]:
"""Every test the console has reported as failed, by name, without duplicates."""
seen = []
for line in log.splitlines():
if "FAILED" not in line:
continue
match = FAILED_LINE.match(line.strip())
if match:
name = f"{match.group(1).split('.')[-1]}.{match.group(2)}"
if name not in seen:
seen.append(name)
return seen


def progress_of(log: str):
"""The most recent (done, total, skipped, failed), or None before any test has run."""
found = PROGRESS.findall(log)
return tuple(int(n) for n in found[-1]) if found else None


def leaked_device_locks() -> int:
"""
How many managed-device slots AGP believes are in use.

Above zero with nothing running means the next run waits forever for a slot and boots no
emulator at all - see AGENTS.md. It is the single most misdiagnosable hang here, because it
produces no output whatsoever.
"""
try:
with open(LOCK_FILE, encoding="utf-8") as handle:
match = re.search(r"MDLockCount\s+(\d+)", handle.read())
return int(match.group(1)) if match else 0
except OSError:
return 0


def verdict_from_xml() -> str | None:
"""
The result as the XML reports it, which is the one to trust.

**Scoped to the newest run's own directory.** AGP keeps ``connected/`` and
``managedDevice/`` side by side and neither is cleared, so summing everything on disk mixes
this run with whatever ran before it - a 22-test run reported as 74 the first time this was
written, and a stale green would have masked a fresh red exactly as readily.

The age is always stated. A verdict that is quietly minutes old is the same failure in a
different coat.

Returns None when no results exist - itself an answer, and a different one from "zero
failures": a run that dies before any test reports writes nothing at all, and reading that as
success is how a red build gets called green.
"""
files = glob.glob(RESULT_XML, recursive=True)
if not files:
return None

newest = max(files, key=os.path.getmtime)
run_dir = os.path.dirname(newest)
files = [f for f in files if os.path.dirname(f) == run_dir]

tests = failures = skipped = 0
named = []
for path in files:
root = ET.parse(path).getroot()
tests += int(root.get("tests", 0))
failures += int(root.get("failures", 0)) + int(root.get("errors", 0))
skipped += int(root.get("skipped", 0))
for case in root.iter("testcase"):
if list(case.iter("failure")) or list(case.iter("error")):
named.append(f"{case.get('classname', '').split('.')[-1]}.{case.get('name')}")

age = (time.time() - os.path.getmtime(newest)) / 60
where = os.path.basename(run_dir) or run_dir
summary = f"{tests} tests, {failures} failed, {skipped} skipped ({where}, {age:.0f} min old)"
return summary if not named else summary + " -> " + ", ".join(named)


def describe_now(log: str) -> list[str]:
"""Everything the log currently says, for --once."""
lines = []
done = progress_of(log)
lines.append(f"progress: {done[0]}/{done[1]}, {done[3]} failed, {done[2]} skipped"
if done else "progress: no test has reported yet")

found = failures_in(log)
lines.extend(f"FAILED {name}" for name in found)
if not found:
lines.append("failures: none reported on the console")

lines.append("terminal: " + ("yes" if TERMINAL.search(log) else "not yet"))
lines.append("xml verdict: " + (verdict_from_xml() or "no results written yet"))

locks = leaked_device_locks()
if locks:
lines.append(f"NOTE MDLockCount is {locks} - if nothing is running, that is the hang")
return lines


def diagnose_stall(log: str, age: float) -> list[str]:
"""
Say which of the two silent hangs this is, because the fixes are unrelated.

Whether any test has reported is what separates them. Nothing at all means the run never got
a device - overwhelmingly a leaked managed-device slot. Tests that ran and then stopped means
the device died under them, which no amount of lock-clearing helps.
"""
done = progress_of(log)
where = f"at {done[0]}/{done[1]}" if done else "before any test ran"
lines = [f"STALLED no output for {age / 60:.0f} min, {where}"]

locks = leaked_device_locks()
if not done and locks:
lines.append(f"STALLED MDLockCount is {locks} with no test output - almost certainly "
f"leaked managed-device slots. ./gradlew --stop && rm {LOCK_FILE}")
elif done:
lines.append("STALLED tests had been running, so suspect the device rather than the "
"lock: adb logcat -b crash")
return lines


def read(path: str) -> str:
try:
with open(path, encoding="utf-8", errors="replace") as handle:
return handle.read()
except OSError:
return ""


def final_report(log: str) -> list[str]:
"""What to say once the run is over. The XML has the last word, not the console."""
lines = []
done = progress_of(log)
if done:
lines.append(f"FINISHED {done[0]}/{done[1]}, {done[3]} failed, {done[2]} skipped")
lines.append("VERDICT " + (verdict_from_xml()
or "NO RESULTS WRITTEN - the run died before any test reported"))
return lines


def watch(args) -> int:
"""Poll until the run reaches a terminal state, printing each new event once."""
reported: set[str] = set()
stalled_at = None
started = time.time()

while True:
log = read(args.logfile)

for name in failures_in(log):
if name not in reported:
reported.add(name)
print(f"FAILED {name}", flush=True)

if TERMINAL.search(log):
for line in final_report(log):
print(line, flush=True)
return 0

# **A run that stops producing output is a result too.** Left unsaid it is
# indistinguishable from a slow test, which is how ten minutes goes by twice.
age = time.time() - os.path.getmtime(args.logfile) if os.path.exists(args.logfile) else 0
if age > args.stall_minutes * 60 and stalled_at != int(age // 60):
stalled_at = int(age // 60)
for line in diagnose_stall(log, age):
print(line, flush=True)

if time.time() - started > 3 * 60 * 60:
print("GIVING UP watched for three hours", flush=True)
return 1

time.sleep(args.interval)


def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("logfile", help="the file the gradle run is redirected to")
parser.add_argument("--once", action="store_true",
help="print the current state and exit, to prove the patterns match")
parser.add_argument("--interval", type=int, default=40, help="seconds between polls")
parser.add_argument("--stall-minutes", type=float, default=8.0,
help="say so if the log stops growing for this long")
args = parser.parse_args()

if args.once:
for line in describe_now(read(args.logfile)):
print(line)
return 0

return watch(args)


if __name__ == "__main__":
sys.exit(main())
27 changes: 15 additions & 12 deletions .github/ISSUE_TEMPLATE/app-bug.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,9 @@ body:
value: |
For the **Android app**.

The most useful thing you can attach is the log. The app can write one for you — see the
box near the bottom, which explains where the button is and how to make it appear.
The most useful thing you can attach is the log. The app writes one for you and takes
the personal details out of it first — see the box near the bottom for where the button
is.

- type: input
id: version
Expand Down Expand Up @@ -66,9 +67,9 @@ body:
attributes:
label: Which exporter made the zip?
description: >-
Only if this is about importing or about missing locations. If the ⋮ menu →
**Information** lists what your tags were imported from, copy it from there; otherwise
whichever version you remember downloading, or leave it blank.
Only if this is about importing or about missing locations. The ⋮ menu →
**Information** prints it under the version, and the app's error page prints it too —
copy it from either.
**Do not open the zip to find out.** It holds the keys to your tags, newer ones are
password-protected on purpose, and unpacking it to read a version line is not worth the
risk of what you might then attach.
Expand All @@ -81,8 +82,9 @@ body:
description: |
There are two ways to get one, and the first is easier if it is offered:

**If the app showed you an error page with an Export logs button**, use that button. It is
the same log, without any of the setting-up below.
**If the app showed you a page saying "This one is a bug"**, use its **Get the log**
button. It offers to copy the log for pasting into the box below, or to save it as a file
to attach — and it needs none of the setting-up below.

**Otherwise** the button is hidden, because most people never need it:

Expand All @@ -97,10 +99,11 @@ body:
keys to your tags: anybody who has it can locate them, and it cannot be un-posted. Nothing
that gets asked here needs it.

**This is a raw Android log and nothing is removed from it.** Read it before posting and
take out anything you would not put on a public page: your Apple ID, device names, serial
numbers, anything a notification happened to say. Unlike the desktop exporter's Save logs
button, this one does not redact.
**Both routes clean the log first**, and tell you what they took out — Apple IDs, device
names, serial numbers, coordinates, place names and bundle passwords, replaced by
placeholders that stay consistent so the log still reads. It is best effort, not a
guarantee: it removes what has a recognisable shape, and something unusual on your
account may not. Worth a glance before you post, not a line-by-line audit.

Never paste your Apple ID password, a verification code, or a device passcode. No answer
requires them.
Expand All @@ -111,7 +114,7 @@ body:
attributes:
label: Before you post
options:
- label: If I attached a log, I have read it and taken out anything personal I did not want public.
- label: If I attached a log, I have had a look at it and taken out anything personal the app missed.
required: true

- type: markdown
Expand Down
Loading
Loading