Skip to content

Remember missing actor state keys instead of re-reading them - #1238

Open
CasperGN wants to merge 6 commits into
dapr:mainfrom
CasperGN:feat/actor-cache-missing-state
Open

CasperGN wants to merge 6 commits into
dapr:mainfrom
CasperGN:feat/actor-cache-missing-state

Conversation

@CasperGN

@CasperGN CasperGN commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Based on the PR for fix/actor-refresh-default-tracker (#1237, which is itself based on #1227). Merge it after that. Until then this diff also shows the earlier commits; this PR's own commits are 02f0066 ("feat(actor): cache state that is not found in the default state tracker") and b491245 ("test(actor): run the nested reentrant test on Python 3.10").

Description

What was wrong. Reading an actor state key that does not exist went to the state store every time, because the state manager did not remember misses. Code that checks for a key and then creates it (contains_state then set_state, get_or_add_state, try_add_state) paid for an extra store read each time.

What changed. When try_get_state finds a key missing, it now stores a "not found" entry in the default state tracker (the one activation, reminders and timers use). Later calls use it with no store read:

  • try_get_state, get_state, contains_state report the key as absent (get_state raises KeyError).
  • try_remove_state / remove_state have nothing to remove (return False / raise KeyError).
  • try_add_state, set_state, set_state_ttl, get_or_add_state, add_or_update_state turn the entry into an add, keeping the ttl. add_or_update_state does not call its update function.
  • get_state_names leaves the key out, is_state_marked_for_remove returns False, and save_state never sends the entry.

After a reentrant save, the default-tracker refresh from #1237 treats a "not found" entry like a clean one: a reentrant write replaces it with the saved value, and a reentrant remove drops it. Without this, a reminder that saw a key as missing would keep seeing it as missing after a reentrant call created it.

Design choices.

  • No new member on the public StateChangeKind enum (the .NET SDK added NotFound). A new member would break user code that matches on every value, and the change-kind-to-operation map in _state_provider.py would raise KeyError if it ever reached a save. The entry is a private StateMetadata subclass that reports the existing none change kind, so any code that does not know about it treats it as clean and never saves it. Public types, ActorStateChange and subclasses of ActorStateManager are unchanged.
  • Misses are cached only in the default tracker, not in the per-call trackers used under reentrancy. Only the default tracker is refreshed after a reentrant save. If an outer reentrant call cached a miss and a nested reentrant call on the same actor (A -> B -> A) then created the key, the outer call would still see it as missing, and get_or_add_state / try_add_state would overwrite the nested call's value. A test covers this case. As far as I can tell, dotnet-sdk#1913 caches misses in the per-call tracker too and may have the same gap.

Behaviour change. Missing keys read through the default tracker now stay cached until the actor deactivates or clear_cache() is called. An actor that reads many distinct missing keys (for example names built from input) keeps one small entry per key for its lifetime. The .NET SDK behaves the same way.

This ports dapr/dotnet-sdk#1913 and its follow-up fix dapr/dotnet-sdk#1916. Thanks to @olitomlinson for pointing out the .NET changes on #1227.

Issue reference

Related to dapr/dapr#10532 (reminders miss state saved by reentrant calls) and #1227.

Checklist

  • Code compiles correctly
  • Created/updated tests
  • Extended the documentation (no public API change)

Ran: pytest tests/actor (205 passed, 10 new), ruff check, ruff format --check, mypy (all clean).

🤖 Generated with Claude Code

JoshVanL and others added 5 commits September 22, 2026 15:10
With reentrancy enabled, each dispatched method call gets its own
state change tracker, but activation, reminders and timers run on the
default tracker because no reentrancy id reaches them. A key read
during activation stays cached there with change kind none forever,
while method calls write the same key through their own trackers.

A reminder callback that later reads that key is served the stale
activation value. An app that skips its write because the value looks
unchanged loses that write silently: nothing is logged anywhere,
because no write is ever issued.

Drop the default tracker's clean copies of keys written through a
reentrancy-scoped tracker, so the next read reloads them from the
runtime.

Reported in dapr/dapr#10532, where a reminder callback's
read-modify-write of an actor state key never persisted while the
identical write from an ordinary method call did, and only with
reentrancy enabled.

Should be backported.

Signed-off-by: joshvanl <me@joshvanl.dev>
…ant save

Instead of dropping the default tracker's clean copy of a key that a
reentrant call saved, replace it with the saved value and ttl so the next
read from activation, a reminder or a timer is served from cache rather
than costing an extra state store read. Removed keys are still dropped and
entries with pending changes are left alone.

The cached value is passed through the state serializer first, so it has
the same shape a fresh read would return (for example a tuple comes back
as a list). This mirrors dapr/dotnet-sdk#1912.

Signed-off-by: Casper Nielsen <casper@diagrid.io>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tch a fresh read

Two cases fall back to dapr#1227's eviction instead of an in-place refresh:
a saved None value, which the state provider leaves out of the write, and
a state serializer that fails to decode the value after the save has
already committed. The save no longer raises after a successful write.

Signed-off-by: Casper Nielsen <casper@diagrid.io>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reading a key that does not exist went back to the state store every time,
because a miss was not cached. try_get_state now records a miss as a clean
"not found" entry in the default tracker, so later reads, contains_state,
try_remove_state and remove_state answer without I/O, and try_add_state,
set_state, set_state_ttl, get_or_add_state and add_or_update_state turn it
into an add without asking the store again. get_state_names skips it and
save_state never sends it.

Misses are not cached in reentrant trackers. Nothing refreshes an outer
reentrant call's tracker when a nested reentrant call (A -> B -> A) saves,
so a cached miss there would hide the nested write and then overwrite it
through get_or_add_state or try_add_state.

The entry is a private StateMetadata subclass that reports the existing
"none" change kind, so the public StateChangeKind enum, ActorStateChange and
subclasses of ActorStateManager are unchanged.

A reentrant save refreshes a "not found" default entry the same way as a
clean one: a write replaces it with the saved value and a remove evicts it.
Without that, a reminder that cached a miss would keep seeing the key as
absent after a reentrant call created it.

Ports dapr/dotnet-sdk#1913 and dapr/dotnet-sdk#1916.

Signed-off-by: Casper Nielsen <casper@diagrid.io>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov

codecov Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 84.01%. Comparing base (fb229bc) to head (b491245).

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #1238      +/-   ##
==========================================
+ Coverage   83.89%   84.01%   +0.11%     
==========================================
  Files         123      123              
  Lines       10265    10301      +36     
==========================================
+ Hits         8612     8654      +42     
+ Misses       1653     1647       -6     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

create_task(context=...) only exists from Python 3.11. A task already runs
in a copy of the caller's context, so the argument isn't needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Casper Nielsen <casper@diagrid.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants