The Selector Service (token/services/selector) picks the unspent tokens (UTXOs) that fund a transaction and holds them under a temporary lock while the transaction is assembled, so that concurrent transactions of the same wallet do not try to spend the same tokens.
The Selector Service is responsible for:
The Selector Service bridges the gap between the high-level TTX Service and the internal TokenDB.
graph LR
TTX[TTX Service] --> Selector[Selector Service]
subgraph "Token Fetcher"
Fetcher[Fetcher Logic]
Fetcher -->|Cache Hit| Cache[Cache]
Fetcher -->|Cache Miss| TokenDB[Token Store - TokenDB]
end
subgraph "Selection Logic"
Query[Query Spendable Tokens]
Pick[Take Next Candidate - randomized order]
Lock[Acquire Temporary Lock]
Done[Return Locked Tokens]
end
Selector --> Fetcher
Selector --> Query
Query --> Pick
Pick --> Lock
Lock -->|locked by another process, or sum still below target| Pick
Lock -->|requested amount covered| Done
How the components interact:
The SelectorManager is the entry point for obtaining a Selector instance anchored to a specific transaction. It ensures that the selection process is consistent and tied to the lifecycle of a single token request.
Selection is a randomized greedy first-fit. It is not configurable, and it is not
amount-aware. Selector.selectInternal (token/services/selector/sherdlock/selector.go)
does the following:
token.SelectorRateLimited is a hard abort
(not a skip),A token’s amount therefore only decides when the loop stops, never which candidate is picked. Two consequences worth planning for:
The randomization is deliberate. It is what spreads concurrent selectors of the same
wallet across different candidates: walking a fixed order would make every selector contend
for the same first tokens, driving up lock failures and, with them, the immediate-retry path
that gives up with token.SelectorSufficientButLockedFunds, and beyond it the backoff path
that ends in token.SelectorInsufficientFunds.
The shuffle lives in the sherdlock fetcher, not in the selection loop
(token/services/selector/sherdlock/fetcher.go): the lazy fetcher wraps the database
iterator in collections.NewPermutatedIterator, and the cached fetcher hands out a fresh
permutation of the cached slice on every query. The simple driver does not shuffle — it
walks the database iterator in the order the token store returns it
(token/services/selector/simple/selector.go) — so concurrent selectors under simple are
more exposed to colliding on the same leading candidates.
How it works in the flow (see “Selection Logic” subgraph in diagram):
TokenLocks table under the sherdlock driver, in memory under simple); on success its amount is added to the running sum, on failure the loop moves to the next candidatesherdlock only): the inner loop refetches — refreshing the
sherdlock token cache via the fetcher — up to a hardcoded maxImmediateRetries = 5 times
without releasing its already-acquired locks, then gives up with
token.SelectorSufficientButLockedFunds. Under simple, there is no equivalent cache
layer; the outer retry loop re-queries the query service directly on every attempt.numRetries / retryInterval outer loop (the
StubbornSelector wrapper in sherdlock; the numRetry / timeout loop in simple)
releases locks, sleeps, and re-runs the whole selection from scratch. Exhausting this
layer returns token.SelectorInsufficientFunds.Amount-aware strategies — smallest-first, largest-first, First-In-First-Out, or minimizing the number of inputs — are not implemented and cannot be configured. There is no strategy abstraction in the code and no configuration key that selects one. Making selection amount-aware is tracked in issue #2017.
To prevent double-spending before the transaction is committed to the ledger, the Selector Service uses a local TokenLocks table in the Storage Service (see “TokenLocks” box in diagram above).
Lock lifecycle:
TokenLocks table.Every leaseCleanupTickPeriod, sherdlock runs a cleanup pass over the TokenLocks
table that releases a lock when either of the following holds.
Both
leaseExpiryandleaseCleanupTickPeriodmust be non-zero for the cleanup goroutine to start. If either is zero the pass never runs, so locks held byDeletedorOrphanconsumers are never released and those tokens remain permanently unselectable. SettingleaseExpiry: 0to disable time-based expiry while relying on consumer-status release is therefore not supported.
consumer_tx_id — has reached Deleted or Orphan, so it will never spend the
token; orleaseExpiry, which covers the consumer that crashed or was
abandoned without ever reaching a terminal status.Two properties of the pass are worth spelling out:
(tx_id, idx), which identifies the locked token and therefore the transaction
that created it. That transaction’s status says nothing about whether the lock is
still live, so it is never used to expire a lease.(tx_id, idx) rows are
deleted; the other outputs of the same transaction keep their locks.The pass is the same statement on every SQL backend (SQLite and Postgres), so lock
expiry behaviour is identical across those backends. created_at is stored as
TIMESTAMPTZ, so the comparison with the database-side NOW() expression is always
timezone-consistent on Postgres regardless of the session TimeZone setting.
On Postgres a single replica per TMS runs the pass per tick, elected through an advisory
lock; SQLite is non-distributed and always runs it locally.
The in-memory locker described below does not use the TokenLocks table and never
expires locks via Cleanup; its lifecycle is entirely managed in process.
The simple driver keeps its locks in memory (token/services/selector/simple/inmemory)
instead of the TokenLocks table. Its state is sharded per owner (the wallet the tokens
are selected for): every owner has its own shard, holding that owner’s locked tokens
behind its own mutex, and the shards themselves live in a registry map behind a second
mutex. Two owners therefore never serialize against each other, not even while a lock
attempt is waiting on a transaction-status lookup.
Two invariants keep the two mutex levels safe:
IsLocked, UnlockByTxID, the background collector, the locked-token count)
therefore snapshots the registry, releases the registry lock, and only then takes the
individual shard locks. Taking the two in the opposite order deadlocks the locker.Lock that had already obtained that shard
re-checks the mark under the shard lock and retries on the freshly registered shard,
so a lock can never end up in a shard no other operation can reach. Pruning also
removes the registry entry only if it still points at that exact shard, so a stale
empty shard cannot evict a newer shard holding live locks.The background collector (the goroutine that frees locks of finalized transactions) copies a shard’s entries, releases the shard lock, and only then looks the transaction statuses up, so a slow status provider never blocks locking or unlocking. Because the shard is unlocked in between, each entry is re-validated before removal — same transaction ID and same last-access time — and entries that were reclaimed or re-accessed meanwhile are kept.
The selector uses a Token Fetcher to retrieve available tokens from the database. The fetcher uses a Ristretto LRU cache to improve performance by caching token queries (keyed by wallet+currency).
Flow: Selector.Select() → Fetcher.UnspentTokensIteratorBy(wallet, currency) → Token Iterator
How it works:
Iterator lifecycle: the iterator the fetcher hands out owns a resource — on the lazy
path it wraps the query’s sql.Rows, and therefore a database cursor and its pooled
connection — so exactly one Close() per iterator is required. The selector holds at most
one iterator at a time: step 6 above closes the iterator it displaces before installing the
refreshed one, and Selector.Close() closes whichever is current. Both happen under the
selector’s mutex, so closing a selector while a retry is in flight neither leaks an iterator
nor races with the retry.
Adaptive refresh strategy with two triggers, which differ in who pays for the refresh:
fetcherCacheRefresh. This refresh is
synchronous — the requesting goroutine runs the query itself and waits for it. An
untouched cache counts as stale, so the first request after startup always pays for it.fetcherCacheMaxQueries queries to prevent serving stale
data in high-throughput scenarios. This refresh runs in the background, and the request
that triggered it is still answered from the current cache.Either way the refresh is a single query for all wallets and currencies
(SpendableTokensIteratorBy with empty arguments), not one per cache key.
fetcherStrategy picks which of the three implemented fetchers the sherdlock driver uses
(token/services/selector/sherdlock/fetcher.go). These are the only accepted values, and an
unrecognized one makes the node fail to start rather than falling back silently:
fetcherStrategy |
Behaviour | Cache keys apply |
|---|---|---|
mixed (default) |
Serves a request from the cache, and falls back to a database query when the cache holds no token for that wallet and currency. | yes |
eager |
Always answers from the cache, never from a query scoped to the requested wallet. The cache itself is refreshed on the two triggers above, so a request that finds it stale does pay for the refreshing query inline. | yes |
lazy |
Always queries the database and keeps no cache at all. | no |
The trade-off is staleness against query load. lazy never offers a token that was already
spent, at the cost of one query per selection request; eager amortizes its queries over many
requests — most are answered from memory, and the ones that refresh a stale cache pay for a
single query covering every wallet — but it can offer tokens spent since the last refresh,
which the selector then fails to lock and retries over. mixed is the default because it
takes the cache’s fast path only when the cache has
something to say about the wallet, and pays for a query otherwise. Which path mixed takes is
observable: the panurus_services_selector_sherdlock_unspent_tokens_invocations counter is
incremented with fetcher_type set to eager or lazy on every request. Only mixed
reports it — under eager and lazy there is no choice to record, and the counter stays at
zero (see Metrics).
The choice is per node, not per TMS: one strategy is resolved at startup and every TMS’s fetcher is built from it.
Configure the selector service in your core.yaml:
token:
selector:
driver: sherdlock # Selector implementation and locking backend: sherdlock | simple (default: sherdlock)
numRetries: 3 # Retry attempts for token selection (default: 3)
retryInterval: 5s # Wait time between retries (default: 5s)
leaseExpiry: 3m # Lock expiration time (default: 3m)
leaseCleanupTickPeriod: 1m # Lock cleanup interval (default: 1m)
fetcherStrategy: mixed # Token fetcher: mixed | eager | lazy (default: mixed)
fetcherCacheSize: 1000 # Cache size in entries (default: 0 = use fetcher default)
fetcherCacheRefresh: 30s # Cache refresh interval (default: 0 = use fetcher default)
fetcherCacheMaxQueries: 100 # Max queries before cache refresh (default: 0 = use fetcher default)
driver selects the selector implementation and, with it, the locking backend:
TokenLocks table of the Storage Service, with
leases governed by leaseExpiry and leaseCleanupTickPeriod.It does not select a selection algorithm: both drivers walk candidates greedily and stop on first cover, but they diverge in several ways beyond the shuffle:
sherdlock randomizes the candidate order; simple walks tokens in database order.sherdlock holds already-acquired locks across immediate retries; simple releases all
locks between every retry attempt.simple runs a GetTokens concurrency check after a successful cover and can return a
fourth error sentinel, token.SelectorSufficientFundsButConcurrencyIssue, which
sherdlock does not produce.mixed (default), eager or
lazy. See Fetcher strategies. Leave it unset to take the default;
any other value than the three above aborts startup.The three keys below tune the cache that mixed and eager use, and are ignored by lazy:
Example: With fetcherCacheSize: 1000, fetcherCacheRefresh: 30s, and fetcherCacheMaxQueries: 100, the cache stores up to 1000 query results, refreshes data every 30 seconds, and forces a refresh after 100 queries.