Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
80 changes: 80 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,86 @@ Each recommendation result includes `cost_ms`, the wall-clock time spent in the
native search itself. Batch results report this value independently for every
request.

## Pool reuse

A `recommend` call spends most of its time building the candidate pool and only
a small fraction searching it. When the same user, masterdata and options are
queried repeatedly, build the pool once and search it many times:

```python
pool = engine.build_pool(options)
result = pool.recommend() # search only
top5 = pool.recommend(limit=5) # limit and timeout_ms may be overridden
print(pool.card_count) # candidates in the pool
```

Measured on one dataset (672-card account, `multi` / `score`, 194 candidates):
`engine.recommend()` 4549 us versus `pool.recommend()` 448 us at the median,
about 10x, saving roughly 4.1 ms per call. Both paths return identical decks.

### What a pool is bound to

A pool captures the user data, the masterdata and the options it was built from.
**It does not observe later changes to any of them**, so reusing a stale pool
silently returns results computed from outdated inputs. Rebuild when:

| Change | Effect |
|---|---|
| `update_masterdata` / `update_musicmetas` | every pool for that region is stale |
| the user's cards change (new cards, levels, master ranks) | that user's pools are stale |
| any option other than `limit` / `timeout_ms` | needs its own pool |

`limit` and `timeout_ms` affect only the search stage and can be passed per call.

### Memory

A pool holds its candidate set, search context and resolved card details until
it is released. Measured per pool:

| Account cards | Candidates | Per pool |
|---|---|---|
| 672 | 141-194 | 221-257 KB |
| 1249 | 156-260 | 391-465 KB |

Roughly 0.2-0.5 MB each, so keeping 100 pools costs about 22-47 MB.

### Caller-owned cache

No pool cache is built in: the right bound and the right invalidation depend on
the caller, and pools cost memory that the library should not claim on its own.
A bounded LRU is a few lines:

```python
from collections import OrderedDict


class PoolCache:
def __init__(self, engine, max_pools=64):
self._engine = engine
self._max = max_pools
self._pools = OrderedDict()

def recommend(self, key, options, limit=None):
pool = self._pools.pop(key, None)
if pool is None:
pool = self._engine.build_pool(options)
self._pools[key] = pool # newest last
while len(self._pools) > self._max:
self._pools.popitem(last=False) # evict oldest
return pool.recommend(limit=limit)

def drop_user(self, user_id):
for key in [k for k in self._pools if k[0] == user_id]:
del self._pools[key]

def clear(self):
self._pools.clear()
```

`key` must cover everything the pool is bound to; a workable one is
`(user_id, user_data_revision, options_fingerprint)`. Call `drop_user` when that
user's cards change and `clear` after reloading masterdata.

## License

MIT
4 changes: 2 additions & 2 deletions python/allium_deck/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,8 @@
DeckRecommendOptions as RecommendOptions,
DeckRecommendResult as RecommendResult,
DeckRecommendUserData as UserData,
PreparedCardPool as CardPool,
SekaiDeckRecommend as Engine,
)

__all__ = ["Engine", "RecommendOptions", "RecommendResult", "UserData"]

__all__ = ["CardPool", "Engine", "RecommendOptions", "RecommendResult", "UserData"]
68 changes: 68 additions & 0 deletions python/sekai_deck_recommend_cpp/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -145,6 +145,28 @@ def recommend(self, options: DeckRecommendOptions) -> DeckRecommendResult:
)
return DeckRecommendResult.from_dict(json.loads(payload))

def build_pool(self, options: DeckRecommendOptions) -> PreparedCardPool:
"""Build a reusable search pool for `options`.

A `recommend` call spends most of its time building the candidate pool
and only a small fraction searching it. When the same user, masterdata
and options are queried repeatedly, build the pool once and search it
many times.

The pool is bound to the user data, masterdata and options it was built
from, and does not observe later changes to any of them. Rebuild it when
any of those change; holding pools costs extra memory. See the pool
reuse section of the README.
"""
region = self._validate_options(options)
user_data = self._resolve_user_data(options)
native_pool = self._require_native().build_pool(
region,
json.dumps(options._to_native_dict(), separators=(",", ":")),
user_data._native,
)
return PreparedCardPool(native_pool)

def recommend_batch(
self, options_list: list[DeckRecommendOptions]
) -> list[DeckRecommendResult]:
Expand Down Expand Up @@ -234,6 +256,51 @@ def calculate_exact_live(
return json.loads(payload)


class PreparedCardPool:
"""A pre-built search pool returned by `SekaiDeckRecommend.build_pool`.

Searching a pool skips pool construction, which is the dominant cost of a
`recommend` call. The pool holds the candidate set, the search context and
the resolved card details, so it occupies memory proportional to the
candidate count until it is released.

It does not track the user data, masterdata or options it was built from.
If any of those change, discard the pool and build a new one.
"""

__slots__ = ("_native",)

def __init__(self, native) -> None:
self._native = native

@property
def card_count(self) -> int:
"""Number of candidate cards in the pool."""
return self._native.card_count

@property
def limit(self) -> int:
"""The `limit` the pool was built with."""
return self._native.limit

@property
def timeout_ms(self) -> int:
"""The `timeout_ms` the pool was built with."""
return self._native.timeout_ms

def recommend(
self, limit: int | None = None, timeout_ms: int | None = None
) -> DeckRecommendResult:
"""Search the pool.

`limit` and `timeout_ms` affect only the search stage and may be
overridden per call. Every other option is fixed at build time; changing
one requires building a new pool.
"""
payload = self._native.recommend(limit, timeout_ms)
return DeckRecommendResult.from_dict(json.loads(payload))


__all__ = [
"DeckRecommendCardConfig",
"DeckRecommendGaOptions",
Expand All @@ -242,6 +309,7 @@ def calculate_exact_live(
"DeckRecommendSaOptions",
"DeckRecommendSingleCardConfig",
"DeckRecommendUserData",
"PreparedCardPool",
"RecommendCard",
"RecommendDeck",
"RecommendSupportDeckCard",
Expand Down
12 changes: 12 additions & 0 deletions python/sekai_deck_recommend_cpp/__init__.pyi
Original file line number Diff line number Diff line change
Expand Up @@ -181,6 +181,7 @@ class SekaiDeckRecommend:
def update_musicmetas(self, file_path: str, region: str) -> None: ...
def update_musicmetas_from_string(self, data: Union[str, bytes], region: str) -> None: ...
def recommend(self, options: DeckRecommendOptions) -> DeckRecommendResult: ...
def build_pool(self, options: DeckRecommendOptions) -> PreparedCardPool: ...
def recommend_batch(
self, options_list: List[DeckRecommendOptions]
) -> List[DeckRecommendResult]: ...
Expand All @@ -204,4 +205,15 @@ class SekaiDeckRecommend:
fever_music_score_json: Optional[str] = None,
) -> Dict[str, Any]: ...

class PreparedCardPool:
@property
def card_count(self) -> int: ...
@property
def limit(self) -> int: ...
@property
def timeout_ms(self) -> int: ...
def recommend(
self, limit: int | None = None, timeout_ms: int | None = None
) -> DeckRecommendResult: ...

def set_engine_thread_count(threads: int) -> None: ...
Loading
Loading