From 7d1866df7b12a36e98f0b2f339df6b9c555a12e5 Mon Sep 17 00:00:00 2001 From: Remy Degenne Date: Sun, 13 Sep 2026 08:55:37 +0200 Subject: [PATCH 1/7] rename doc --- .../SequentialLearning/Algorithms/Markov.lean | 0 .../SequentialLearning/README.md | 157 ++++++++++++++++++ 2 files changed, 157 insertions(+) create mode 100644 LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean create mode 100644 LeanMachineLearning/SequentialLearning/README.md diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean b/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean new file mode 100644 index 00000000..e69de29b diff --git a/LeanMachineLearning/SequentialLearning/README.md b/LeanMachineLearning/SequentialLearning/README.md new file mode 100644 index 00000000..a2fa4af1 --- /dev/null +++ b/LeanMachineLearning/SequentialLearning/README.md @@ -0,0 +1,157 @@ +# Naming of algorithms and environments + +The file `Algorithm.lean` defines the `Algorithm` and `Environment` structures. +An algorithm has a sequence of Markov kernels (the policy) which reads a history and an observation and returns an action. +```lean +structure Algorithm (π“ž 𝓐 𝓨 : Type*) [MeasurableSpace π“ž] [MeasurableSpace 𝓐] [MeasurableSpace 𝓨] + where + /-- Law of the action of round `n` given the past rounds and the current observation. -/ + policy : (n : β„•) β†’ Kernel (Hist π“ž 𝓐 𝓨 n Γ— π“ž) 𝓐 + /-- The policy is a Markov kernel. -/ + [isMarkovKernel_policy : βˆ€ n, IsMarkovKernel (policy n)] +``` +An environment has two sequences of kernels: one for observations, which reads the history, and one for feedback, which reads the history, the current observation and the current action. +```lean +structure Environment (π“ž 𝓐 𝓨 : Type*) [MeasurableSpace π“ž] [MeasurableSpace 𝓐] [MeasurableSpace 𝓨] + where + /-- Law of the observation of round `n` given the past rounds. -/ + obs : (n : β„•) β†’ Kernel (Hist π“ž 𝓐 𝓨 n) π“ž + /-- Law of the feedback of round `n` given the past rounds, the observation and the action. -/ + feedback : (n : β„•) β†’ Kernel ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐) 𝓨 + /-- The observation kernel is a Markov kernel. -/ + [isMarkovKernel_obs : βˆ€ n, IsMarkovKernel (obs n)] + /-- The feedback kernel is a Markov kernel. -/ + [isMarkovKernel_feedback : βˆ€ n, IsMarkovKernel (feedback n)] +``` + +In many applications, some of those kernels are deterministic, or do not depend on some of their inputs. +We detail here the naming conventions for the various constructors, predicates and accessors that are used in the library. + +Generic constructions live in the `Algorithm` and `Environment` namespaces. +A `det` prefix marks the deterministic version of a constructor. +Predicates are root-level `Is…Alg` / `Is…Env` classes when they carry an accessor, and namespaced `Prop` definitions otherwise. +Accessors are namespaced so that dot notation works. +Every row may read the time `n` unless it says "stationary" or "nothing". +Every `det…` constructor is the stochastic one applied to `Kernel.deterministic`, gets both the determinism instance and the dependence instance, and takes its measurability proofs as `by fun_prop` autoparams. + +All time zero accessors are root-level `…0` definitions. Example: `Algorithm.policy0`. + +## Algorithms + +The policy at round `n` can depend on `n`, on the history at `n` and the current observation. It can be deterministic or stochastic. + +Not stochastic: `Algorithm.IsDeterministic`, `Algorithm.deterministic` + +No observation: `Algorithm.IgnoresObs`, `Algorithm.comapObs fun _ ↦ ()` + +No history: `Algorithm.IsMarkov`, `Algorithm.markov` + +No history, no observation: `Algorithm.IsOpenLoop`, `Algorithm.openLoop`, `Algorithm.ofSeq` (det version) + +Not time-dependent, no history: `Algorithm.IsStationary`, `Algorithm.stationary` + +No time, no history, no observation: `Algorithm.const` + +## Environments + +The observation at round `n` can depend on `n` and on the history at `n`, and can be deterministic or stochastic. + +The feedback at round `n` can depend on `n`, on the history at `n`, on the current observation and on the current action. +It can be deterministic or stochastic. + +In general, the dependence on history is the same for both kernels. + +All for obs, no action for feedback: `Environment.FeedbackIgnoresAction`, `Environment.adversary`. + +No history for obs and feedback: `Environment.IsOblivious`, `Environment.oblivious`. + +No time, no history for obs and feedback: `Environment.IsStationary`, `Environment.stationary`. + +No time, last round of history for obs, not history for feedback: `Environment.IsMarkov`, `Environment.markov`. + +Determinism: `Environment.HasDeterministicObs`, `Environment.HasDeterministicFeedback`. + +### Obs = Unit + +Only the feedback matters, so the constructors are named by the feedback's shape. +It can depend on time, history and action and be stochastic or deterministic. + +general: use `Environment` with Obs = Unit. + +No history: use oblivious or stationary environment with Obs = Unit. + +No history, deterministic: `Environment.evalSeq` and `Environment.eval` (no time). + +No history, no action (only time): `Environment.indep` and `Environment.ofSeq` (deterministic). + +Nothing: `Environment.const` (stochastic). The deterministic version is probably not useful. + +## Examples + +Oblivious adversarial bandit environment: Obs = Unit, feedback depends on time and action, deterministic. Which constructor? `Environment.evalSeq`. + +Stochastic optimization: Obs = Unit, feedback depends on time and action, stochastic. Which constructor? `Environment.banditSeq`. + +| Policy reads | Stochastic constructor | Deterministic constructor | Predicate | Accessor | +|---|---|---|---|---| +| history, obs | `Algorithm` | `Algorithm.deterministic f`, was `detAlgorithm` | `IsDeterministicAlg`, exists | `alg.nextAction n`, was `nextAction alg n` | +| history only | `alg.comapObs fun _ ↦ ()`, exists | same | `Algorithm.IgnoresObs`, new def | none | +| obs | `Algorithm.markov Ο€`, `Ο€ : β„• β†’ Kernel π“ž 𝓐`, new | `Algorithm.detMarkov f`, `f : β„• β†’ π“ž β†’ 𝓐`, new | `IsMarkovAlg`, new | `alg.policyCondObs n : Kernel π“ž 𝓐` | +| obs, stationary | `Algorithm.stationary Ο€`, `Ο€ : Kernel π“ž 𝓐`, new | `Algorithm.detStationary f`, `f : π“ž β†’ 𝓐`, new | `IsStationaryAlg`, new | same, constant in `n` | +| time only | `Algorithm.openLoop ΞΌ`, `ΞΌ : β„• β†’ Measure 𝓐`, new | `Algorithm.ofSeq x`, was `fixedDesignAlg` | `IsOpenLoopAlg`, new | `alg.actionLaw n : Measure 𝓐` | +| nothing | `Algorithm.const ΞΌ`, was `randomSampling` | `Algorithm.detConst a`, new | none, use `IsOpenLoopAlg` | none | + +Named instances of rows: `Algorithm.uniform`, was `uniformAlgorithm`, is `const` of the uniform +measure; `Algorithm.roundRobin hK`, was `roundRobinAlgorithm`, is `ofSeq`. The bandit-specific +`ucbAlgorithm`, `etcAlgorithm` and `tsAlgorithm` keep their names in the `Bandits` namespace. + +Implications provided as instances: stationary implies Markov, open loop implies Markov, and the +deterministic constructors of each row give both instances. The three dependence predicates are +the named cases of `Algorithm.FactorsThrough`. + +## Environments + +### General observation type + +| Obs reads | Feedback reads | Stochastic constructor | Deterministic constructor | Predicate | Accessors | +|---|---|---|---|---|---| +| history | history, obs, action | the structure | `Environment.deterministic g f`, new | `IsDeterministicEnv`, redefined to cover both kernels | `env.obsFun n`, `env.feedbackFun n` | +| history | history, obs | `Environment.adversary ρ ΞΊ`, new | `Environment.detAdversary g f`, new | `Environment.FeedbackIgnoresAction`, from LMLPapers | none | +| last round | obs, action | `Environment.markov ρ P Ξ½`, new | `Environment.detMarkov sβ‚€ P f`, new | `IsMarkovEnv`, new | `env.transition n : Kernel (Round π“ž 𝓐 𝓨) π“ž`, `env.feedbackCondObsAction n : Kernel (π“ž Γ— 𝓐) 𝓨` | +| last obs and action, stationary | obs, action | `Environment.mdp ρ P r`, new | `Environment.detMdp sβ‚€ P r`, new | none, use `IsMarkovEnv` | same | +| time only | obs, action | `Environment.oblivious ρ Ξ½`, `ρ : β„• β†’ Measure π“ž`, `Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨`, new signature | `Environment.detOblivious o f`, new | `IsObliviousEnv`, redefined | `env.obsLaw n : Measure π“ž`, `env.feedbackCondObsAction n` | +| nothing | obs, action, stationary | `Environment.stationary ρ Ξ½`, `ρ : Measure π“ž`, `Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨`, new signature | `Environment.detStationary oβ‚€ f`, new | `IsStationaryEnv`, new | same, constant in `n` | + +Today's `detEnvironment obs f`, with random observations and deterministic feedback, becomes +`Environment.detFeedback obs f` with the definition `Environment.HasDeterministicFeedback`, or is +dropped since nothing uses it. + +### No observation + +With `π“ž = Unit` only the feedback matters, so these constructors are named by the feedback's shape +rather than by a dependence word. They are the existing bandit constructors, unchanged in type. + +| Feedback reads | Stochastic constructor | Deterministic constructor | +|---|---|---| +| time, action | `Environment.banditSeq Ξ½`, `Ξ½ : β„• β†’ Kernel 𝓐 𝓨`, was `obliviousEnv` | `Environment.evalSeq g`, was `onlineEvalEnv` | +| action | `Environment.bandit Ξ½`, `Ξ½ : Kernel 𝓐 𝓨`, was `stationaryEnv` | `Environment.eval f`, was `evalEnv` | +| time only | `Environment.indep P`, `P : β„• β†’ Measure 𝓨`, new | `Environment.ofSeq y`, was `seqEnv` in LMLPapers | +| nothing | `Environment.iid P`, was `iidEnv` in LMLPapers | `Environment.const y`, new | + +The general predicates apply to this table as they are. The one extra accessor is +`env.feedbackCondAction n : Kernel 𝓐 𝓨`, defined only for `π“ž = Unit` from +`feedbackCondObsAction`; it is today's `feedbackCondAction env n`. The identities tying the two +tables together are simp lemmas: `bandit Ξ½` is `stationary` of the Dirac measure and +`Ξ½.prodMkLeft Unit`, `eval f` is `detStationary`, `ofSeq y` is `evalSeq` of constant functions, +`iid P` is `bandit` of the constant kernel, and `indep P` is `banditSeq` of constant kernels. + +### Hidden parameter + +`Environment.bayes Q env : Environment (𝓔 Γ— π“ž) 𝓐 𝓨` for a family `env : 𝓔 β†’ Environment π“ž 𝓐 𝓨`, +new; `Environment.bayesBandit Q ΞΊ : Environment 𝓔 𝓐 𝓨`, was `bayesStationaryEnv`. + +Implications provided as instances: stationary implies oblivious implies Markov, and `mdp` is a +`markov` instance. + +Time zero: `Environment.obs0` unchanged; `Environment.feedback0`, was `Ξ½0`; +`Environment.feedbackFun0`, was `feedbackFunZero`. From 34883b41b9aec4658a10af8e50945018ab89df50 Mon Sep 17 00:00:00 2001 From: Remy Degenne Date: Thu, 17 Sep 2026 10:08:48 +0200 Subject: [PATCH 2/7] generalize oblivious and stationary env --- LMLTutorial/Pages/DefiningAlgorithm.lean | 14 +- .../Online/Bandit/Algorithms/ETC.lean | 24 +- .../Online/Bandit/Algorithms/Regret/ETC.lean | 8 +- .../Online/Bandit/Algorithms/Regret/UCB.lean | 28 +- .../Online/Bandit/Algorithms/TS.lean | 4 +- .../Online/Bandit/Algorithms/UCB.lean | 19 +- .../Online/Bandit/ArrayProbSpace.lean | 9 +- .../Online/Bandit/BayesRegret.lean | 2 +- .../Online/Bandit/RewardByCountMeasure.lean | 58 +- .../Online/Bandit/SumRewards.lean | 36 +- .../SequentialLearning/Algorithm.lean | 16 +- .../Algorithms/RoundRobin.lean | 14 +- .../BayesStationaryEnv.lean | 17 +- .../DivergenceDecomposition.lean | 20 +- .../SequentialLearning/EvaluationEnv.lean | 31 +- .../SequentialLearning/Means.lean | 21 +- .../SequentialLearning/README.md | 73 +-- .../SequentialLearning/StationaryEnv.lean | 515 +++++++++++++----- 18 files changed, 556 insertions(+), 353 deletions(-) diff --git a/LMLTutorial/Pages/DefiningAlgorithm.lean b/LMLTutorial/Pages/DefiningAlgorithm.lean index 585cfa43..4940f870 100644 --- a/LMLTutorial/Pages/DefiningAlgorithm.lean +++ b/LMLTutorial/Pages/DefiningAlgorithm.lean @@ -73,17 +73,19 @@ The `Environment` structure is the mirror of the `Algorithm` structure, with a k The distribution of the first observation is `obs 0` applied to the empty history; it is called `Environment.obs0`. The distribution of the first feedback given the first observation and action is `feedback 0` applied to the empty history; it is called `Environment.Ξ½0`. -In many applications there is no observation and the feedback depends only on the last action, not on the prior history. -We provide an `obliviousEnv` definition that builds an environment for those cases. +In many applications neither the observation nor the feedback depends on the prior history: the observation at time `n` has a fixed law, and the feedback depends only on the current observation and action. +We provide an `obliviousEnv` definition that builds an environment for those cases from a sequence of observation laws and a sequence of feedback kernels. {docstring obliviousEnv} -`(Ξ½ n).prodMkLeft _` is the kernel `Ξ½ n` seen as a `Kernel ((Hist Unit 𝓐 𝓨 n Γ— Unit) Γ— 𝓐) 𝓨` by ignoring the history and the observation. - -If furthermore the feedback kernel does not change with time, we can use the `stationaryEnv` definition to build the environment. +If furthermore those sequences do not change with time, we can use the `stationaryEnv` definition to build the environment. {docstring stationaryEnv} +When there is no observation (`π“ž = Unit`), the feedback depends only on the last action. `Environment.banditSeq` builds such an environment from a sequence of kernels `Ξ½ : β„• β†’ Kernel 𝓐 𝓨`, and `Environment.bandit` from a single kernel `Ξ½ : Kernel 𝓐 𝓨` used at every time. + +{docstring Environment.bandit} + # Sequences of actions and feedback, probability space @@ -111,7 +113,7 @@ We now illustrate the use of `Algorithm`, `Environment`, and `IsAlgEnvSeq` by de In a stochastic bandit, an algorithm chooses at each time an action from a finite set (here `Fin K`, the type of natural numbers less than `K`) and receives a reward drawn from a distribution that depends only on the action, not on the prior history. -The environment is thus simply `stationaryEnv Ξ½` for some kernel `Ξ½ : Kernel (Fin K) ℝ`. +The environment is thus simply `Environment.bandit Ξ½` for some kernel `Ξ½ : Kernel (Fin K) ℝ`. ## Algorithm diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean b/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean index 12c60a34..3e014b73 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean @@ -61,8 +61,8 @@ variable [NeZero K] {m : β„•} {Ξ½ : Kernel (Fin K) ℝ} [IsMarkovKernel Ξ½] /-- Before round `K * m`, the ETC algorithm behaves like the Round-Robin algorithm. -/ lemma isAlgEnvSeqUntil_roundRobinAlgorithm - (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) : - IsAlgEnvSeqUntil O A R (roundRobinAlgorithm K) (stationaryEnv Ξ½) P (K * m) := by + (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) : + IsAlgEnvSeqUntil O A R (roundRobinAlgorithm K) (Environment.bandit Ξ½) P (K * m) := by refine h.isAlgEnvSeqUntil_of_policy_eq fun n hn ↦ ?_ simp only [roundRobinAlgorithm, detAlgorithm_policy, etcAlgorithm] congr 1 with p @@ -70,12 +70,13 @@ lemma isAlgEnvSeqUntil_roundRobinAlgorithm section AlgorithmBehavior -lemma arm_ae_eq_nextArm (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) (n : β„•) : +lemma arm_ae_eq_nextArm (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) + (n : β„•) : A n =ᡐ[P] fun Ο‰ ↦ nextArm K m n (history O A R n Ο‰) := h.action_detAlgorithm_ae_eq n /-- For `n < K * m`, the arm pulled at time `n` is the arm `n % K`. -/ -lemma arm_of_lt (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +lemma arm_of_lt (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) {n : β„•} (hn : n < K * m) : A n =ᡐ[P] fun _ ↦ RoundRobin.nextAction K n := RoundRobin.action_ae_eq n ((isAlgEnvSeqUntil_roundRobinAlgorithm h).mono hn) @@ -83,13 +84,13 @@ lemma arm_of_lt (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) /-- The arm pulled at time `K * m` is the arm with the highest empirical mean after the exploration phase. -/ lemma arm_mul - (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) : + (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) : A (K * m) =ᡐ[P] fun Ο‰ ↦ argmax (empMean' (K * m) (history O A R (K * m) Ο‰)) := by filter_upwards [arm_ae_eq_nextArm h (K * m)] with Ο‰ hn_eq rw [hn_eq, nextArm, dite_eq_right (by simp), dite_eq_left rfl] /-- For `n β‰₯ K * m`, the arm pulled at time `n + 1` is the same as the arm pulled at time `n`. -/ -lemma arm_add_one_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +lemma arm_add_one_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) {n : β„•} (hn : K * m ≀ n) : A (n + 1) =ᡐ[P] fun Ο‰ ↦ A n Ο‰ := by filter_upwards [arm_ae_eq_nextArm h (n + 1)] with Ο‰ hn_eq @@ -97,7 +98,7 @@ lemma arm_add_one_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv rfl /-- For `n β‰₯ K * m`, the arm pulled at time `n` is the same as the arm pulled at time `K * m`. -/ -lemma arm_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +lemma arm_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) {n : β„•} (hn : K * m ≀ n) : A n =ᡐ[P] A (K * m) := by have h_ae n : K * m ≀ n β†’ A (n + 1) =ᡐ[P] fun Ο‰ ↦ A n Ο‰ := arm_add_one_of_ge h @@ -108,11 +109,12 @@ lemma arm_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) | succ n hmn h_ind => rw [h_ae n hmn, h_ind] /-- At time `K * m`, the number of pulls of each arm is equal to `m`. -/ -lemma pullCount_mul (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) (a : Fin K) : +lemma pullCount_mul (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) + (a : Fin K) : pullCount A a (K * m) =ᡐ[P] fun _ ↦ m := RoundRobin.pullCount_mul m (isAlgEnvSeqUntil_roundRobinAlgorithm h) a -lemma pullCount_add_one_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +lemma pullCount_add_one_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (a : Fin K) {n : β„•} (hn : K * m ≀ n) : pullCount A a (n + 1) =ᡐ[P] fun Ο‰ ↦ pullCount A a n Ο‰ + {Ο‰' | A (K * m) Ο‰' = a}.indicator 1 Ο‰ := by @@ -122,7 +124,7 @@ lemma pullCount_add_one_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (station /-- For `n β‰₯ K * m`, the number of pulls of each arm `a` at time `n` is equal to `m` plus `n - K * m` if arm `a` is the best arm after the exploration phase. -/ -lemma pullCount_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +lemma pullCount_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (a : Fin K) {n : β„•} (hn : K * m ≀ n) : pullCount A a n =ᡐ[P] fun Ο‰ ↦ m + (n - K * m) * {Ο‰' | A (K * m) Ο‰' = a}.indicator 1 Ο‰ := by @@ -142,7 +144,7 @@ lemma pullCount_of_ge (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv /-- If at time `K * m` the algorithm chooses arm `a`, then the total reward obtained by pulling arm `a` is at least the total reward obtained by pulling the best arm. -/ lemma sumRewards_bestArm_le_of_arm_mul_eq - (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) (a : Fin K) (hm : m β‰  0) : + (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (a : Fin K) (hm : m β‰  0) : βˆ€α΅ Ο‰ βˆ‚P, A (K * m) Ο‰ = a β†’ sumRewards A R (bestArm Ξ½) (K * m) Ο‰ ≀ sumRewards A R a (K * m) Ο‰ := by filter_upwards [arm_mul h, pullCount_mul h a, pullCount_mul h (bestArm Ξ½)] diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/Regret/ETC.lean b/LeanMachineLearning/Online/Bandit/Algorithms/Regret/ETC.lean index cfc3f1ee..c5a40cfc 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/Regret/ETC.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/Regret/ETC.lean @@ -26,7 +26,7 @@ variable {K : β„•} [NeZero K] {m : β„•} {Ξ½ : Kernel (Fin K) ℝ} [IsMarkovKerne {Οƒ2 : ℝβ‰₯0} lemma probReal_sumRewards_le_sumRewards_le - (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (a : Fin K) : P.real {Ο‰ | sumRewards A R (bestArm Ξ½) (K * m) Ο‰ ≀ sumRewards A R a (K * m) Ο‰} ≀ Real.exp (-↑m * gap Ξ½ a ^ 2 / (4 * Οƒ2)) := by @@ -42,7 +42,7 @@ lemma probReal_sumRewards_le_sumRewards_le /-- The probability that at time `K * m` the ETC algorithm chooses arm `a` is at most `exp(- m * Ξ”_a^2 / (4 * Οƒ2))`. -/ -lemma prob_arm_mul_eq_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +lemma prob_arm_mul_eq_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (a : Fin K) (hm : m β‰  0) : P.real {Ο‰ | A (K * m) Ο‰ = a} ≀ Real.exp (- (m : ℝ) * gap Ξ½ a ^ 2 / (4 * Οƒ2)) := by @@ -57,7 +57,7 @@ lemma prob_arm_mul_eq_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEn exact h_le.trans (probReal_sumRewards_le_sumRewards_le h hΞ½ a) /-- Bound on the expectation of the number of pulls of each arm by the ETC algorithm. -/ -lemma expectation_pullCount_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +lemma expectation_pullCount_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (a : Fin K) (hm : m β‰  0) {n : β„•} (hn : K * m ≀ n) : P[fun Ο‰ ↦ (pullCount A a n Ο‰ : ℝ)] @@ -85,7 +85,7 @@ lemma expectation_pullCount_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (statio Β· exact (measurableSet_singleton _).preimage (by fun_prop) /-- Regret bound for the ETC algorithm. -/ -theorem regret_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (stationaryEnv Ξ½) P) +theorem regret_le (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hm : m β‰  0) (n : β„•) (hn : K * m ≀ n) : P[regret Ξ½ A n] ≀ diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/Regret/UCB.lean b/LeanMachineLearning/Online/Bandit/Algorithms/Regret/UCB.lean index 7d106788..275e3dba 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/Regret/UCB.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/Regret/UCB.lean @@ -75,7 +75,7 @@ lemma pullCount_le_of_ucbIndex_le (hc : 0 ≀ c) {b : Fin K} /-- The probability that the UCB index of arm `a` is below its mean is at most `1 / (n + 1) ^ (c - 1)`. -/ lemma prob_ucbIndex_le {alg : Algorithm Unit (Fin K) ℝ} - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 0 ≀ c) (a : Fin K) (n : β„•) : P {Ο‰ | 0 < pullCount A a n Ο‰ ∧ empMean A R a n Ο‰ + ucbWidth A (c * Οƒ2) a n Ο‰ ≀ (Ξ½ a)[id]} ≀ @@ -102,7 +102,7 @@ lemma prob_ucbIndex_le {alg : Algorithm Unit (Fin K) ℝ} /-- The probability that the lower confidence bound of arm `a` is above its mean is at most `1 / (n + 1) ^ (c - 1)`. -/ lemma prob_lcbIndex_ge {alg : Algorithm Unit (Fin K) ℝ} - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 0 ≀ c) (a : Fin K) (n : β„•) : P {Ο‰ | 0 < pullCount A a n Ο‰ ∧ @@ -170,7 +170,7 @@ lemma pullCount_le_add_three (a : Fin K) (n C : β„•) (Ο‰ : Ξ©) : rw [Finset.sum_add_distrib, Finset.sum_add_distrib] lemma pullCount_le_add_three_ae - (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (a : Fin K) (n C : β„•) (hC : C β‰  0) : βˆ€α΅ Ο‰ βˆ‚P, pullCount A a n Ο‰ ≀ C + 1 + @@ -195,7 +195,7 @@ lemma pullCount_le_add_three_ae at which it already has more than `C` pulls and the means of the best arm and of `a` lie in their confidence intervals. -/ lemma sum_indicator_good_event_eq_zero - (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (Environment.bandit Ξ½) P) (hc : 0 ≀ c) (a : Fin K) (h_gap : 0 < gap Ξ½ a) (n C : β„•) (hC : C β‰  0) (hC' : 8 * c * Οƒ2 * log (n + 1) / gap Ξ½ a ^ 2 ≀ C) : βˆ€α΅ Ο‰ βˆ‚P, @@ -228,7 +228,8 @@ lemma sum_indicator_good_event_eq_zero simp_rw [← mul_assoc] gcongr -lemma pullCount_ae_le_add_two (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (stationaryEnv Ξ½) P) +lemma pullCount_ae_le_add_two + (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (Environment.bandit Ξ½) P) (hc : 0 ≀ c) (a : Fin K) (h_gap : 0 < gap Ξ½ a) (n C : β„•) (hC : C β‰  0) (hC' : 8 * c * Οƒ2 * log (n + 1) / gap Ξ½ a ^ 2 ≀ C) : βˆ€α΅ Ο‰ βˆ‚P, @@ -296,7 +297,7 @@ lemma constSum_le {c : ℝ} (hc : 2 < c) (n : β„•) : constSum c n ≀ 1 + 1 / (c /-- Bound on the expectation of the number of pulls of each arm by the UCB algorithm. -/ lemma expectation_pullCount_le' - (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 0 < c) (a : Fin K) (h_gap : 0 < gap Ξ½ a) (n : β„•) : ∫⁻ Ο‰, pullCount A a n Ο‰ βˆ‚P ≀ @@ -381,7 +382,8 @@ lemma expectation_pullCount_le' positivity /-- Bound on the expectation of the number of pulls of each arm by the UCB algorithm. -/ -lemma expectation_pullCount_le (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (stationaryEnv Ξ½) P) +lemma expectation_pullCount_le + (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 0 < c) (a : Fin K) (h_gap : 0 < gap Ξ½ a) (n : β„•) : P[fun Ο‰ ↦ (pullCount A a n Ο‰ : ℝ)] ≀ @@ -403,7 +405,7 @@ lemma expectation_pullCount_le (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) ring /-- Regret bound for the UCB algorithm. -/ -lemma regret_le (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (stationaryEnv Ξ½) P) +lemma regret_le (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 0 < c) (n : β„•) : P[regret Ξ½ A n] ≀ @@ -417,7 +419,7 @@ lemma regret_le (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (stationaryEnv Β· field /-- Regret bound for the UCB algorithm with an explicit constant, for `c > 2`. -/ -lemma regret_le_of_gt_two (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (stationaryEnv Ξ½) P) +lemma regret_le_of_gt_two (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 2 < c) (n : β„•) : P[regret Ξ½ A n] ≀ @@ -431,26 +433,26 @@ lemma regret_le_of_gt_two (h : IsAlgEnvSeq O A R (ucbAlgorithm K (c * Οƒ2)) (sta /-- Regret bound for the UCB algorithm with exploration constant `c`, for `Οƒ2`-subgaussian rewards. -/ -lemma regret_le' (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) +lemma regret_le' (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 0 < c) (n : β„•) : P[regret Ξ½ A n] ≀ βˆ‘ a, (8 * c * log (n + 1) / gap Ξ½ a + gap Ξ½ a * (2 + 2 * constSum (c / Οƒ2) n)) := by have hΟƒ2' : (0 : ℝ) < Οƒ2 := NNReal.coe_pos.mpr (pos_iff_ne_zero.mpr hΟƒ2) - have h' : IsAlgEnvSeq O A R (ucbAlgorithm K (c / Οƒ2 * Οƒ2)) (stationaryEnv Ξ½) P := by + have h' : IsAlgEnvSeq O A R (ucbAlgorithm K (c / Οƒ2 * Οƒ2)) (Environment.bandit Ξ½) P := by rwa [div_mul_cancelβ‚€ _ hΟƒ2'.ne'] refine (regret_le h' hΞ½ hΟƒ2 (div_pos hc hΟƒ2') n).trans_eq ?_ rw [show (8 : ℝ) * (c / Οƒ2) * Οƒ2 = 8 * c by rw [mul_assoc, div_mul_cancelβ‚€ _ hΟƒ2'.ne']] /-- Regret bound for the UCB algorithm with exploration constant `c > 2 * Οƒ2`, for `Οƒ2`-subgaussian rewards, with an explicit constant. -/ -theorem regret_le_of_gt_two' (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) +theorem regret_le_of_gt_two' (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) (hΟƒ2 : Οƒ2 β‰  0) (hc : 2 * Οƒ2 < c) (n : β„•) : P[regret Ξ½ A n] ≀ βˆ‘ a, (8 * c * log (n + 1) / gap Ξ½ a + gap Ξ½ a * (4 + 2 * Οƒ2 / (c - 2 * Οƒ2))) := by have hΟƒ2' : (0 : ℝ) < Οƒ2 := NNReal.coe_pos.mpr (pos_iff_ne_zero.mpr hΟƒ2) - have h' : IsAlgEnvSeq O A R (ucbAlgorithm K (c / Οƒ2 * Οƒ2)) (stationaryEnv Ξ½) P := by + have h' : IsAlgEnvSeq O A R (ucbAlgorithm K (c / Οƒ2 * Οƒ2)) (Environment.bandit Ξ½) P := by rwa [div_mul_cancelβ‚€ _ hΟƒ2'.ne'] have hc' : 2 < c / Οƒ2 := by rwa [lt_div_iffβ‚€ hΟƒ2'] refine (regret_le_of_gt_two h' hΞ½ hΟƒ2 hc' n).trans_eq ?_ diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean b/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean index aa380f73..4ffe5cb1 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean @@ -20,7 +20,7 @@ probability of being optimal under the posterior over environments given the his * `tsAlgorithm Q ΞΊ`: a Thompson sampling algorithm with actions in `Fin K` (for `K β‰  0`), given a prior distribution over parameters `Q : Measure 𝓔` and a Markov kernel `ΞΊ : Kernel (𝓔 Γ— Fin K) ℝ`. This kernel defines how a parameter `e : 𝓔` gives rise to - a stationary environment: `stationaryEnv (ΞΊ.sectR e) : Environment (Fin K) ℝ`. + a stationary environment: `Environment.bandit (ΞΊ.sectR e) : Environment (Fin K) ℝ`. ## Main results @@ -60,7 +60,7 @@ instance [NeZero K] {Q : Measure 𝓔} [IsProbabilityMeasure Q] {ΞΊ : Kernel ( /-- The Thompson sampling algorithm with actions in `Fin K`, where `Q : Measure 𝓔` is a prior distribution over parameters, and `ΞΊ : Kernel (𝓔 Γ— Fin K) ℝ` is a Markov kernel that defines the - stationary environment `stationaryEnv (ΞΊ.sectR e)` that corresponds to a parameter `e : 𝓔`. + stationary environment `Environment.bandit (ΞΊ.sectR e)` that corresponds to a parameter `e : 𝓔`. At every time `n`, the Thompson sampling policy uses the posterior over the parameters given the history up to time `n` to derive the probability of each action being optimal. The action for time diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean b/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean index 6522921b..6efa3880 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean @@ -66,8 +66,8 @@ variable [NeZero K] {c : ℝ} {Ξ½ : Kernel (Fin K) ℝ} [IsMarkovKernel Ξ½] /-- Before round `K`, the UCB algorithm behaves like the Round-Robin algorithm. -/ lemma isAlgEnvSeqUntil_roundRobinAlgorithm - (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) : - IsAlgEnvSeqUntil O A R (roundRobinAlgorithm K) (stationaryEnv Ξ½) P K := by + (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) : + IsAlgEnvSeqUntil O A R (roundRobinAlgorithm K) (Environment.bandit Ξ½) P K := by refine h.isAlgEnvSeqUntil_of_policy_eq fun n hn ↦ ?_ simp only [roundRobinAlgorithm, detAlgorithm_policy, ucbAlgorithm] congr 1 with p @@ -92,16 +92,17 @@ lemma ucbWidth_eq_ucbWidth' (c : ℝ) (a : Fin K) (n : β„•) (Ο‰ : Ξ©) : ucbWidth A c a n Ο‰ = ucbWidth' c n (history O A R n Ο‰) a := by rw [ucbWidth, ucbWidth', pullCount_eq_pullCount' (O := O) (A := A) (R' := R)] -lemma arm_zero (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) : +lemma arm_zero (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) : A 0 =ᡐ[P] fun _ ↦ 0 := RoundRobin.action_zero ((isAlgEnvSeqUntil_roundRobinAlgorithm h).mono (Nat.pos_of_neZero K)) -lemma arm_ae_eq_nextArm (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) (n : β„•) : +lemma arm_ae_eq_nextArm (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) + (n : β„•) : A n =ᡐ[P] fun Ο‰ ↦ nextArm K c n (history O A R n Ο‰) := h.action_detAlgorithm_ae_eq n lemma ucbIndex_le_ucbIndex_arm - (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) (a : Fin K) (hn : K ≀ n) : + (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (a : Fin K) (hn : K ≀ n) : βˆ€α΅ Ο‰ βˆ‚P, empMean A R a n Ο‰ + ucbWidth A c a n Ο‰ ≀ empMean A R (A n Ο‰) n Ο‰ + ucbWidth A c (A n Ο‰) n Ο‰ := by filter_upwards [arm_ae_eq_nextArm h n] with Ο‰ h_arm @@ -111,7 +112,7 @@ lemma ucbIndex_le_ucbIndex_arm exact isMaxOn_argmax (fun a ↦ empMean' n (history O A R n Ο‰) a + ucbWidth' c n (history O A R n Ο‰) a) _ -lemma forall_arm_eq_mod_of_lt (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) : +lemma forall_arm_eq_mod_of_lt (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n < K, A n Ο‰ = RoundRobin.nextAction K n := by simp_rw [ae_all_iff] intro n hn @@ -120,7 +121,7 @@ lemma forall_arm_eq_mod_of_lt (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (station simp only [nextArm, hn, ↓reduceIte] lemma forall_ucbIndex_le_ucbIndex_arm - (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) (a : Fin K) : + (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (a : Fin K) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, K ≀ n β†’ empMean A R a n Ο‰ + ucbWidth A c a n Ο‰ ≀ empMean A R (A n Ο‰) n Ο‰ + ucbWidth A c (A n Ο‰) n Ο‰ := by @@ -128,12 +129,12 @@ lemma forall_ucbIndex_le_ucbIndex_arm exact fun _ ↦ ucbIndex_le_ucbIndex_arm h a lemma time_gt_of_pullCount_gt_one - (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) (a : Fin K) : + (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (a : Fin K) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, 1 < pullCount A a n Ο‰ β†’ K < n := RoundRobin.time_gt_of_pullCount_gt_one (isAlgEnvSeqUntil_roundRobinAlgorithm h) a lemma pullCount_pos_of_pullCount_gt_one - (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (stationaryEnv Ξ½) P) (a : Fin K) : + (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (a : Fin K) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, 1 < pullCount A a n Ο‰ β†’ βˆ€ b : Fin K, 0 < pullCount A b n Ο‰ := RoundRobin.pullCount_pos_of_pullCount_gt_one (isAlgEnvSeqUntil_roundRobinAlgorithm h) a diff --git a/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean b/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean index 3bbfc3a6..b5320351 100644 --- a/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean +++ b/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean @@ -755,7 +755,7 @@ lemma hasCondDistrib_reward (alg : Algorithm Unit 𝓐 𝓑) (Ξ½ : Kernel 𝓐 (n : β„•) : HasCondDistrib (reward alg n) (fun Ο‰ ↦ ((history (noObs _) (action alg) (reward alg) n Ο‰, noObs _ n Ο‰), action alg n Ο‰)) - ((stationaryEnv Ξ½).feedback n) (arrayMeasure Ξ½) := by + ((Environment.bandit Ξ½).feedback n) (arrayMeasure Ξ½) := by let e : (Hist Unit 𝓐 𝓑 n Γ— 𝓐) ≃ᡐ ((Hist Unit 𝓐 𝓑 n Γ— Unit) Γ— 𝓐) := { toFun := fun p ↦ ((p.1, ()), p.2) invFun := fun p ↦ (p.1.1, p.2) @@ -763,13 +763,14 @@ lemma hasCondDistrib_reward (alg : Algorithm Unit 𝓐 𝓑) (Ξ½ : Kernel 𝓐 right_inv := fun _ ↦ rfl measurable_toFun := by simp only [Equiv.coe_fn_mk]; fun_prop measurable_invFun := by simp only [Equiv.symm_mk, Equiv.coe_fn_mk]; fun_prop } - rw [feedback_stationaryEnv] + rw [feedback_bandit] have h := (hasCondDistrib_reward' alg Ξ½ n).measurableEquiv_comp_right e simp only [hist_eq_history] at h exact h lemma isAlgEnvSeq_arrayMeasure (alg : Algorithm Unit 𝓐 𝓑) (Ξ½ : Kernel 𝓐 𝓑) [IsMarkovKernel Ξ½] : - IsAlgEnvSeq (noObs _) (action alg) (reward alg) alg (stationaryEnv Ξ½) (arrayMeasure Ξ½) where + IsAlgEnvSeq (noObs _) (action alg) (reward alg) alg (Environment.bandit Ξ½) + (arrayMeasure Ξ½) where hasCondDistrib_obs n := hasCondDistrib_unit (measurable_history (fun _ ↦ measurable_const) (measurable_action alg) (measurable_reward alg) n).aemeasurable _ _ @@ -786,7 +787,7 @@ lemma hasCondDistrib_reward_zero (alg : Algorithm Unit 𝓐 𝓑) (Ξ½ : Kernel [IsMarkovKernel Ξ½] : HasCondDistrib (reward alg 0) (action alg 0) Ξ½ (arrayMeasure Ξ½) := by have h := (isAlgEnvSeq_arrayMeasure alg Ξ½).hasCondDistrib_feedback_zero - rw [Ξ½0_stationaryEnv] at h + rw [Ξ½0_bandit] at h simpa using hasCondDistrib_prodMk_left_unique_iff.mp h end Laws diff --git a/LeanMachineLearning/Online/Bandit/BayesRegret.lean b/LeanMachineLearning/Online/Bandit/BayesRegret.lean index 00210b41..c1a95147 100644 --- a/LeanMachineLearning/Online/Bandit/BayesRegret.lean +++ b/LeanMachineLearning/Online/Bandit/BayesRegret.lean @@ -17,7 +17,7 @@ measurable space `Ξ©`. These definitions are useful when `IsBayesAlgEnvSeq Q ΞΊ Recall that `IsBayesAlgEnvSeq Q ΞΊ alg E A Y P` states that there is a measure `P : Measure Ξ©` such that the parameter `E : Ξ© β†’ 𝓔` has law `Q` and that the sequences of actions `A : β„• β†’ Ξ© β†’ 𝓐` and feedbacks `Y : β„• β†’ Ξ© β†’ 𝓨` are generated by the algorithm `alg : Algorithm 𝓐 𝓨` interacting with an -underlying environment that depends on `E` and `ΞΊ` (`stationaryEnv (ΞΊ.sectR (E Ο‰))`) +underlying environment that depends on `E` and `ΞΊ` (`Environment.bandit (ΞΊ.sectR (E Ο‰))`) ## Main definitions diff --git a/LeanMachineLearning/Online/Bandit/RewardByCountMeasure.lean b/LeanMachineLearning/Online/Bandit/RewardByCountMeasure.lean index fe804ee9..29e24b5a 100644 --- a/LeanMachineLearning/Online/Bandit/RewardByCountMeasure.lean +++ b/LeanMachineLearning/Online/Bandit/RewardByCountMeasure.lean @@ -22,7 +22,7 @@ namespace Bandits variable {𝓐 Ξ© : Type*} {m𝓐 : MeasurableSpace 𝓐} {mΞ© : MeasurableSpace Ξ©} [DecidableEq 𝓐] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {R : β„• β†’ Ξ© β†’ ℝ} {P : Measure Ξ©} [IsProbabilityMeasure P] {alg : Algorithm Unit 𝓐 ℝ} {Ξ½ : Kernel 𝓐 ℝ} [IsMarkovKernel Ξ½] - {h_inter : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P} + {h_inter : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P} local notation "𝔓" => P.prod (streamMeasure Ξ½) @@ -35,11 +35,11 @@ notation "𝓛[" Y " | " X " ← " x "; " ΞΌ "]" => Measure.map Y (ΞΌ[|X ⁻¹' omit [DecidableEq 𝓐] in lemma condDistrib_reward'' [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (n : β„•) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (n : β„•) : 𝓛[fun Ο‰ ↦ R n Ο‰.1 | fun Ο‰ ↦ A n Ο‰.1; 𝔓] =ᡐ[(𝔓).map (fun Ο‰ ↦ A n Ο‰.1)] Ξ½ := by have hA := h.measurable_action have hR := h.measurable_feedback - have h_ra' : 𝓛[R n | A n; P] =ᡐ[P.map (A n)] Ξ½ := h.condDistrib_feedback_stationaryEnv n + have h_ra' : 𝓛[R n | A n; P] =ᡐ[P.map (A n)] Ξ½ := h.condDistrib_feedback_bandit n have h_law : (𝔓).map (fun Ο‰ ↦ A n Ο‰.1) = P.map (A n) := by change ((𝔓).map (A n ∘ Prod.fst)) = _ rw [← Measure.map_map (by fun_prop) (by fun_prop), ← Measure.fst, Measure.fst_prod] @@ -56,7 +56,7 @@ variable [StandardBorelSpace 𝓐] omit [DecidableEq 𝓐] in lemma reward_cond_action [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (n : β„•) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (n : β„•) (hΞΌa : (𝔓).map (fun Ο‰ ↦ A n Ο‰.1) {a} β‰  0) : 𝓛[fun Ο‰ ↦ R n Ο‰.1 | fun Ο‰ ↦ A n Ο‰.1 ← a; 𝔓] = Ξ½ a := by have hA := h.measurable_action @@ -74,7 +74,7 @@ lemma reward_cond_action [Countable 𝓐] variable [Nonempty 𝓐] lemma condIndepFun_reward_stepsUntil_action' [StandardBorelSpace Ξ©] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (m n : β„•) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (m n : β„•) : R n βŸ‚α΅’[A n, h.measurable_action n; P] {Ο‰ | stepsUntil A a m Ο‰ = ↑n}.indicator (fun _ ↦ 1) := by -- the indicator of `stepsUntil ... = n` is a function of `hist (n-1)` and `action n`. -- It thus suffices to use the independence of `reward n` and `hist (n-1)` conditionally @@ -83,12 +83,12 @@ lemma condIndepFun_reward_stepsUntil_action' [StandardBorelSpace Ξ©] have hR := h.measurable_feedback have h_indep : R n βŸ‚α΅’[A n, hA n; P] fun Ο‰ ↦ ((history O A R n Ο‰, O n Ο‰), A n Ο‰) := - IsAlgEnvSeq.condIndepFun_feedback_history_action_action h n + IsAlgEnvSeq.condIndepFun_feedback_history_action_action_bandit h n refine h_indep.of_measurable_right (hX := hA n) ?_ exact measurable_comap_indicator_stepsUntil_eq O R a m n lemma condIndepFun_reward_stepsUntil_action [StandardBorelSpace Ξ©] [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (m n : β„•) : CondIndepFun (m𝓐.comap (fun Ο‰ ↦ A n Ο‰.1)) ((h.measurable_action n).comp measurable_fst).comap_le (fun Ο‰ ↦ R n Ο‰.1) ({Ο‰ | stepsUntil A a m Ο‰.1 = ↑n}.indicator (fun _ ↦ 1)) 𝔓 := by @@ -99,7 +99,7 @@ lemma condIndepFun_reward_stepsUntil_action [StandardBorelSpace Ξ©] [Countable (condIndepFun_reward_stepsUntil_action' h a m n) lemma reward_cond_stepsUntil [StandardBorelSpace Ξ©] [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (m n : β„•) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (m n : β„•) (hm : m β‰  0) (hΞΌn : 𝔓 ((fun Ο‰ ↦ stepsUntil A a m Ο‰.1) ⁻¹' {↑n}) β‰  0) : 𝓛[fun Ο‰ ↦ R n Ο‰.1 | fun Ο‰ ↦ stepsUntil A a m Ο‰.1 ← ↑n; 𝔓] = Ξ½ a := by have hA := h.measurable_action @@ -145,7 +145,7 @@ lemma reward_cond_stepsUntil [StandardBorelSpace Ξ©] [Countable 𝓐] /-- The conditional distribution of the reward received at the `m`-th pull of action `a` given the time at which number of pulls is `m` is the constant kernel with value `Ξ½ a`. -/ lemma condDistrib_rewardByCount_stepsUntil [StandardBorelSpace Ξ©] [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (m : β„•) (hm : m β‰  0) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (m : β„•) (hm : m β‰  0) : condDistrib (rewardByCount A R a m) (fun Ο‰ ↦ stepsUntil A a m Ο‰.1) 𝔓 =ᡐ[(𝔓).map (fun Ο‰ ↦ stepsUntil A a m Ο‰.1)] Kernel.const _ (Ξ½ a) := by have hA := h.measurable_action @@ -178,7 +178,7 @@ lemma condDistrib_rewardByCount_stepsUntil [StandardBorelSpace Ξ©] [Countable /-- The reward received at the `m`-th pull of action `a` has law `Ξ½ a`. -/ lemma hasLaw_rewardByCount [StandardBorelSpace Ξ©] [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (m : β„•) (hm : m β‰  0) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (m : β„•) (hm : m β‰  0) : HasLaw (rewardByCount A R a m) (Ξ½ a) 𝔓 where aemeasurable := (measurable_rewardByCount h.measurable_action h.measurable_feedback a m).aemeasurable @@ -198,7 +198,7 @@ lemma hasLaw_rewardByCount [StandardBorelSpace Ξ©] [Countable 𝓐] _ = Ξ½ a := by simp lemma identDistrib_rewardByCount [StandardBorelSpace Ξ©] [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (n m : β„•) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (n m : β„•) (hn : n β‰  0) (hm : m β‰  0) : IdentDistrib (rewardByCount A R a n) (rewardByCount A R a m) 𝔓 𝔓 where aemeasurable_fst := @@ -208,7 +208,7 @@ lemma identDistrib_rewardByCount [StandardBorelSpace Ξ©] [Countable 𝓐] map_eq := by rw [(hasLaw_rewardByCount h a n hn).map_eq, (hasLaw_rewardByCount h a m hm).map_eq] lemma identDistrib_rewardByCount_id [StandardBorelSpace Ξ©] [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (n : β„•) (hn : n β‰  0) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (n : β„•) (hn : n β‰  0) : IdentDistrib (rewardByCount A R a n) id 𝔓 (Ξ½ a) where aemeasurable_fst := (measurable_rewardByCount h.measurable_action h.measurable_feedback a n).aemeasurable @@ -216,7 +216,7 @@ lemma identDistrib_rewardByCount_id [StandardBorelSpace Ξ©] [Countable 𝓐] map_eq := by rw [(hasLaw_rewardByCount h a n hn).map_eq, Measure.map_id] lemma identDistrib_rewardByCount_eval [StandardBorelSpace Ξ©] [Countable 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (n m : β„•) (hn : n β‰  0) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (n m : β„•) (hn : n β‰  0) : IdentDistrib (rewardByCount A R a n) (fun Ο‰ ↦ Ο‰ m a) 𝔓 (streamMeasure Ξ½) := (identDistrib_rewardByCount_id h a n hn).trans (identDistrib_eval_eval_id_streamMeasure Ξ½ m a).symm @@ -282,24 +282,25 @@ lemma indepFun_update_rewardByCountUntil_eval [Countable 𝓐] (hA : βˆ€ n, Meas /-- Conditionally on the event that the action at time `n` is `b` and that `b` was pulled `k` times before, the reward at time `n` is independent of the history before time `n` and of the action at time `n`. -/ -lemma indepFun_history_reward_cond (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) +lemma indepFun_history_reward_cond (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (n : β„•) (b : 𝓐) (k : β„•) : (fun x ↦ ((history O A R n x, O n x), A n x)) βŸ‚α΅’[P[|{x | A n x = b ∧ pullCount A b n x = k}]] R n := by rw [setOf_action_eq_and_pullCount_eq_eq_preimage (O := O) (R' := R)] - exact h.indepFun_history_action_feedback_cond_stationaryEnv n + exact h.indepFun_history_action_feedback_cond_bandit n (measurableSet_snd_eq_and_pullCount'_eq n b k) fun u hu ↦ hu.1 /-- Conditionally on the event that the action at time `t` is `b` and that `b` was pulled `k` times before, the reward at time `t` has law `Ξ½ b`. -/ -lemma hasLaw_reward_cond (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (t : β„•) (b : 𝓐) (k : β„•) +lemma hasLaw_reward_cond (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (t : β„•) (b : 𝓐) + (k : β„•) (hP : P {x | A t x = b ∧ pullCount A b t x = k} β‰  0) : HasLaw (R t) (Ξ½ b) (P[|{x | A t x = b ∧ pullCount A b t x = k}]) := by rw [setOf_action_eq_and_pullCount_eq_eq_preimage (O := O) (R' := R)] at hP ⊒ - exact h.hasLaw_feedback_cond_stationaryEnv t (measurableSet_snd_eq_and_pullCount'_eq t b k) + exact h.hasLaw_feedback_cond_bandit t (measurableSet_snd_eq_and_pullCount'_eq t b k) (fun u hu ↦ hu.1) hP -lemma hasLaw_reward_cond_prod (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (t : β„•) (b : 𝓐) +lemma hasLaw_reward_cond_prod (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (t : β„•) (b : 𝓐) (k : β„•) (hP : P {x | A t x = b ∧ pullCount A b t x = k} β‰  0) : HasLaw (fun Ο‰ ↦ R t Ο‰.1) (Ξ½ b) ((P[|{x | A t x = b ∧ pullCount A b t x = k}]).prod (streamMeasure Ξ½)) := @@ -310,7 +311,7 @@ variable [Countable 𝓐] /-- Conditionally on the event that the action at time `t` is `b` and that `b` was pulled `k` times before, the array `rewardByCountUntil A R t` with the entry `(b, k)` erased is independent of the reward at time `t`. -/ -lemma indepFun_update_rewardByCountUntil_reward (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) +lemma indepFun_update_rewardByCountUntil_reward (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (t : β„•) (b : 𝓐) (k : β„•) : (fun Ο‰ ↦ Function.update (rewardByCountUntil A R t Ο‰) (b, k) 0) βŸ‚α΅’[(P[|{x | A t x = b ∧ pullCount A b t x = k}]).prod (streamMeasure Ξ½)] @@ -343,7 +344,8 @@ lemma indepFun_update_rewardByCountUntil_reward (h : IsAlgEnvSeq O A R alg (stat times before, the arrays `rewardByCountUntil A R (t + 1)` and `rewardByCountUntil A R t` have the same law: they differ only in the entry `(b, k)`, which is `R t` in the first and an auxiliary reward in the second, and both are independent of the rest of the array with law `Ξ½ b`. -/ -lemma identDistrib_rewardByCountUntil_add_one_cond (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) +lemma identDistrib_rewardByCountUntil_add_one_cond + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (t : β„•) (b : 𝓐) (k : β„•) : IdentDistrib (rewardByCountUntil A R (t + 1)) (rewardByCountUntil A R t) ((P[|{x | A t x = b ∧ pullCount A b t x = k}]).prod (streamMeasure Ξ½)) @@ -399,7 +401,7 @@ lemma identDistrib_rewardByCountUntil_add_one_cond (h : IsAlgEnvSeq O A R alg (s (IdentDistrib.of_ae_eq (measurable_rewardByCountUntil hA hR _).aemeasurable h2).symm /-- The law of `rewardByCountUntil A R t` under `𝔓` does not depend on `t`. -/ -lemma identDistrib_rewardByCountUntil_add_one (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) +lemma identDistrib_rewardByCountUntil_add_one (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (t : β„•) : IdentDistrib (rewardByCountUntil A R (t + 1)) (rewardByCountUntil A R t) 𝔓 𝔓 := by have hA := h.measurable_action @@ -424,7 +426,7 @@ lemma identDistrib_rewardByCountUntil_add_one (h : IsAlgEnvSeq O A R alg (statio exact identDistrib_rewardByCountUntil_add_one_cond h t p.1 p.2 /-- The law of `rewardByCountUntil A R t` under `𝔓` is `⨂ (a, m), Ξ½ a`, for all `t`. -/ -lemma hasLaw_rewardByCountUntil (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (t : β„•) : +lemma hasLaw_rewardByCountUntil (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (t : β„•) : HasLaw (rewardByCountUntil A R t) (Measure.infinitePi fun p : 𝓐 Γ— β„• ↦ Ξ½ p.1) 𝔓 := by induction t with | zero => exact hasLaw_rewardByCountUntil_zero P @@ -432,7 +434,7 @@ lemma hasLaw_rewardByCountUntil (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) /-- The array of rewards by count `(a, m) ↦ rewardByCount A R a (m + 1)` has law `⨂ (a, m), Ξ½ a`: its entries are independent, and the entry `(a, m)` has law `Ξ½ a`. -/ -lemma hasLaw_rewardByCount_infinitePi (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) : +lemma hasLaw_rewardByCount_infinitePi (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) : HasLaw (fun Ο‰ (p : 𝓐 Γ— β„•) ↦ rewardByCount A R p.1 (p.2 + 1) Ο‰) (Measure.infinitePi fun p : 𝓐 Γ— β„• ↦ Ξ½ p.1) 𝔓 := by have hY : Measurable fun Ο‰ (p : 𝓐 Γ— β„•) ↦ rewardByCount A R p.1 (p.2 + 1) Ο‰ := @@ -444,14 +446,14 @@ lemma hasLaw_rewardByCount_infinitePi (h : IsAlgEnvSeq O A R alg (stationaryEnv (hasLaw_rewardByCountUntil h) eventually_rewardByCountUntil_eq /-- The reward received at the `(m + 1)`-th pull of action `a` has law `Ξ½ a`. -/ -lemma hasLaw_rewardByCount_add_one (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) +lemma hasLaw_rewardByCount_add_one (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (m : β„•) : HasLaw (rewardByCount A R a (m + 1)) (Ξ½ a) 𝔓 := (hasLaw_eval_infinitePi (fun p : 𝓐 Γ— β„• ↦ Ξ½ p.1) (a, m)).comp (hasLaw_rewardByCount_infinitePi h) /-- The rewards by count `rewardByCount A R a (m + 1)` are independent over all actions `a` and all counts `m`. -/ -lemma iIndepFun_rewardByCount_add_one (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) : +lemma iIndepFun_rewardByCount_add_one (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) : iIndepFun (fun (p : 𝓐 Γ— β„•) Ο‰ ↦ rewardByCount A R p.1 (p.2 + 1) Ο‰) 𝔓 := (iIndepFun_iff_hasLaw_Pi_infinitePi (X := fun (p : 𝓐 Γ— β„•) Ο‰ ↦ rewardByCount A R p.1 (p.2 + 1) Ο‰) (ΞΌ := fun p : 𝓐 Γ— β„• ↦ Ξ½ p.1) @@ -460,7 +462,7 @@ lemma iIndepFun_rewardByCount_add_one (h : IsAlgEnvSeq O A R alg (stationaryEnv /-- The rewards by count `rewardByCount A R a m` for `m β‰  0` are independent over all actions `a` and all counts `m`. -/ -lemma iIndepFun_rewardByCount (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) : +lemma iIndepFun_rewardByCount (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) : iIndepFun (fun (p : {p : 𝓐 Γ— β„• // p.2 β‰  0}) Ο‰ ↦ rewardByCount A R p.1.1 p.1.2 Ο‰) 𝔓 := by have h_eq : (fun (p : {p : 𝓐 Γ— β„• // p.2 β‰  0}) Ο‰ ↦ rewardByCount A R p.1.1 p.1.2 Ο‰) = fun p Ο‰ ↦ rewardByCount A R p.1.1 (p.1.2 - 1 + 1) Ο‰ := by @@ -474,14 +476,14 @@ lemma iIndepFun_rewardByCount (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) : /-- For each action `a`, the rewards by count `(rewardByCount A R a (m + 1))_m` are independent (and by `hasLaw_rewardByCount_add_one` identically distributed with law `Ξ½ a`). -/ -lemma iIndepFun_rewardByCount_add_one_action (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) +lemma iIndepFun_rewardByCount_add_one_action (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) : iIndepFun (fun m Ο‰ ↦ rewardByCount A R a (m + 1) Ο‰) 𝔓 := (iIndepFun_rewardByCount_add_one h).precomp (g := fun m ↦ (a, m)) fun _ _ hmn ↦ (Prod.mk.inj hmn).2 /-- Two distinct rewards by count are independent. -/ -lemma indepFun_rewardByCount (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) +lemma indepFun_rewardByCount (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {a b : 𝓐} {m n : β„•} (hm : m β‰  0) (hn : n β‰  0) (hne : (a, m) β‰  (b, n)) : rewardByCount A R a m βŸ‚α΅’[𝔓] rewardByCount A R b n := (iIndepFun_rewardByCount h).indepFun (i := ⟨(a, m), hm⟩) (j := ⟨(b, n), hn⟩) diff --git a/LeanMachineLearning/Online/Bandit/SumRewards.lean b/LeanMachineLearning/Online/Bandit/SumRewards.lean index 0263c221..1e9221dd 100644 --- a/LeanMachineLearning/Online/Bandit/SumRewards.lean +++ b/LeanMachineLearning/Online/Bandit/SumRewards.lean @@ -157,8 +157,8 @@ lemma pullCount_eq_comp : -- todo: write those lemmas with IdentDistrib instead of equality of maps lemma _root_.Learning.IsAlgEnvSeq.law_sumRewards_unique [MeasurableSingletonClass 𝓐] - (h1 : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) - (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (stationaryEnv Ξ½) P') : + (h1 : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) + (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (Environment.bandit Ξ½) P') : P.map (sumRewards A R a n) = P'.map (sumRewards Aβ‚‚ Rβ‚‚ a n) := by have hA := h1.measurable_action have hR := h1.measurable_feedback @@ -175,8 +175,8 @@ lemma _root_.Learning.IsAlgEnvSeq.law_sumRewards_unique [MeasurableSingletonClas Β· fun_prop lemma _root_.Learning.IsAlgEnvSeq.law_pullCount_sumRewards_unique' [MeasurableSingletonClass 𝓐] - (h1 : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) - (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (stationaryEnv Ξ½) P') : + (h1 : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) + (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (Environment.bandit Ξ½) P') : IdentDistrib (fun Ο‰ a ↦ (pullCount A a n Ο‰, sumRewards A R a n Ο‰)) (fun Ο‰ a ↦ (pullCount Aβ‚‚ a n Ο‰, sumRewards Aβ‚‚ Rβ‚‚ a n Ο‰)) P P' := by have hO := h1.measurable_obs @@ -222,15 +222,15 @@ lemma _root_.Learning.IsAlgEnvSeq.law_pullCount_sumRewards_unique' [MeasurableSi Β· fun_prop lemma _root_.Learning.IsAlgEnvSeq.law_pullCount_sumRewards_unique [MeasurableSingletonClass 𝓐] - (h1 : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) - (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (stationaryEnv Ξ½) P') : + (h1 : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) + (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (Environment.bandit Ξ½) P') : P.map (fun Ο‰ ↦ (pullCount A a n Ο‰, sumRewards A R a n Ο‰)) = P'.map (fun Ο‰ ↦ (pullCount Aβ‚‚ a n Ο‰, sumRewards Aβ‚‚ Rβ‚‚ a n Ο‰)) := ((h1.law_pullCount_sumRewards_unique' h2 (n := n)).comp (u := fun f ↦ f a) (by fun_prop)).map_eq lemma _root_.Learning.IsAlgEnvSeq.identDistrib_pullCount_sumRewards [MeasurableSingletonClass 𝓐] - (h1 : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) - (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (stationaryEnv Ξ½) P') : + (h1 : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) + (h2 : IsAlgEnvSeq Oβ‚‚ Aβ‚‚ Rβ‚‚ alg (Environment.bandit Ξ½) P') : IdentDistrib (fun Ο‰ n a ↦ (pullCount A a n Ο‰, sumRewards A R a n Ο‰)) (fun Ο‰' n a ↦ (pullCount Aβ‚‚ a n Ο‰', sumRewards Aβ‚‚ Rβ‚‚ a n Ο‰')) P P' := by let f (Ο„ : β„• β†’ Round Unit 𝓐 ℝ) (n : β„•) (a : 𝓐) : β„• Γ— ℝ := @@ -263,7 +263,7 @@ variable [Nonempty 𝓐] -- this is what we will use for UCB lemma prob_pullCount_prod_sumRewards_mem_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {s : Set (β„• Γ— ℝ)} [DecidablePred (Β· ∈ Prod.fst '' s)] (hs : MeasurableSet s) : P {Ο‰ | (pullCount A a n Ο‰, sumRewards A R a n Ο‰) ∈ s} ≀ βˆ‘ k ∈ (range (n + 1)).filter (Β· ∈ Prod.fst '' s), @@ -289,7 +289,7 @@ property `p` holds for the number of pulls and the sum of rewards of action `a` least one pull, is at most `n` times a uniform bound on the probability of that property for the sums of `k ∈ [1, n]` i.i.d. rewards. -/ lemma prob_pullCount_pos_and_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (n : β„•) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (n : β„•) {p : β„• β†’ ℝ β†’ Prop} (hp : Measurable fun q : β„• Γ— ℝ ↦ p q.1 q.2) {B : ℝβ‰₯0∞} (hB : βˆ€ k, k β‰  0 β†’ streamMeasure Ξ½ {Ο‰ | p k (βˆ‘ i ∈ range k, Ο‰ i a)} ≀ B) : P {Ο‰ | 0 < pullCount A a n Ο‰ ∧ p (pullCount A a n Ο‰) (sumRewards A R a n Ο‰)} ≀ n * B := by @@ -313,7 +313,7 @@ lemma prob_pullCount_pos_and_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] _ = n * B := by simp lemma prob_pullCount_mem_and_sumRewards_mem_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {s : Set β„•} [DecidablePred (Β· ∈ s)] (hs : MeasurableSet s) {B : Set ℝ} (hB : MeasurableSet B) : P {Ο‰ | pullCount A a n Ο‰ ∈ s ∧ sumRewards A R a n Ο‰ ∈ B} ≀ βˆ‘ k ∈ (range (n + 1)).filter (Β· ∈ s), @@ -332,7 +332,7 @@ lemma prob_pullCount_mem_and_sumRewards_mem_le [Countable 𝓐] [MeasurableSingl simp [hk.2.1] lemma prob_sumRewards_mem_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {B : Set ℝ} (hB : MeasurableSet B) : P (sumRewards A R a n ⁻¹' B) ≀ βˆ‘ k ∈ range (n + 1), streamMeasure Ξ½ {Ο‰ | βˆ‘ i ∈ range k, Ο‰ i a ∈ B} := by @@ -343,7 +343,7 @@ lemma prob_sumRewards_mem_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] rfl lemma prob_pullCount_eq_and_sumRewards_mem_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {m : β„•} (hm : m ≀ n) {B : Set ℝ} (hB : MeasurableSet B) : P {Ο‰ | pullCount A a n Ο‰ = m ∧ sumRewards A R a n Ο‰ ∈ B} ≀ streamMeasure Ξ½ {Ο‰ | βˆ‘ i ∈ range m, Ο‰ i a ∈ B} := by @@ -352,7 +352,7 @@ lemma prob_pullCount_eq_and_sumRewards_mem_le [Countable 𝓐] [MeasurableSingle simpa [hm'] using h_le lemma prob_exists_pullCount_eq_and_sumRewards_mem_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) (a : 𝓐) (m : β„•) {B : Set ℝ} + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (m : β„•) {B : Set ℝ} (hB : MeasurableSet B) : P {Ο‰ | βˆƒ n, pullCount A a n Ο‰ = m ∧ sumRewards A R a n Ο‰ ∈ B} ≀ streamMeasure Ξ½ {Ο‰ | βˆ‘ i ∈ range m, Ο‰ i a ∈ B} := @@ -369,7 +369,7 @@ lemma prob_exists_pullCount_eq_and_sumRewards_mem_le [Countable 𝓐] [Measurabl _ ≀ _ := ArrayModel.prob_exists_pullCount_eq_and_sumRewards_mem_le a m hB lemma probReal_sumRewards_le_sumRewards_le [Fintype 𝓐] [MeasurableSingletonClass 𝓐] - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) (a : 𝓐) (n m₁ mβ‚‚ : β„•) : P.real {Ο‰ | pullCount A (bestArm Ξ½) n Ο‰ = m₁ ∧ pullCount A a n Ο‰ = mβ‚‚ ∧ sumRewards A R (bestArm Ξ½) n Ο‰ ≀ sumRewards A R a n Ο‰} ≀ @@ -468,7 +468,7 @@ end StreamMeasure lemma prob_sumRewards_sub_pullCount_mul_ge_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] {Οƒ2 : ℝβ‰₯0} (hΟƒ2 : 0 < Οƒ2) (ha : HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) {Ξ΄ : ℝ} (hΞ΄ : 0 < Ξ΄) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {Ξ΄ : ℝ} (hΞ΄ : 0 < Ξ΄) : P {Ο‰ | βˆƒ t < n, pullCount A a t Ο‰ β‰  0 ∧ √(2 * pullCount A a t Ο‰ * Οƒ2 * Real.log (1 / Ξ΄)) ≀ sumRewards A R a t Ο‰ - pullCount A a t Ο‰ * (Ξ½ a)[id]} ≀ ENNReal.ofReal ((n - 1) * Ξ΄) := let B (m : β„•) := {x : ℝ | √(2 * m * Οƒ2 * Real.log (1 / Ξ΄)) ≀ x - m * (Ξ½ a)[id]} @@ -502,7 +502,7 @@ lemma prob_sumRewards_sub_pullCount_mul_ge_le [Countable 𝓐] [MeasurableSingle lemma prob_sumRewards_sub_pullCount_mul_le_le [Countable 𝓐] [MeasurableSingletonClass 𝓐] {Οƒ2 : ℝβ‰₯0} (hΟƒ2 : 0 < Οƒ2) (ha : HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) {Ξ΄ : ℝ} (hΞ΄ : 0 < Ξ΄) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {Ξ΄ : ℝ} (hΞ΄ : 0 < Ξ΄) : P {Ο‰ | βˆƒ t < n, pullCount A a t Ο‰ β‰  0 ∧ sumRewards A R a t Ο‰ - pullCount A a t Ο‰ * (Ξ½ a)[id] ≀ -√(2 * pullCount A a t Ο‰ * Οƒ2 * Real.log (1 / Ξ΄))} ≀ ENNReal.ofReal ((n - 1) * Ξ΄) := @@ -537,7 +537,7 @@ lemma prob_sumRewards_sub_pullCount_mul_le_le [Countable 𝓐] [MeasurableSingle lemma prob_sumRewards_sub_pullCount_mul_ge_le_of_Fintype [Fintype 𝓐] [MeasurableSingletonClass 𝓐] {Οƒ2 : ℝβ‰₯0} (hΟƒ2 : 0 < Οƒ2) (hΞ½ : βˆ€ a, HasSubgaussianMGF (fun x ↦ x - (Ξ½ a)[id]) Οƒ2 (Ξ½ a)) - (h : IsAlgEnvSeq O A R alg (stationaryEnv Ξ½) P) {Ξ΄ : ℝ} (hΞ΄ : 0 < Ξ΄) : + (h : IsAlgEnvSeq O A R alg (Environment.bandit Ξ½) P) {Ξ΄ : ℝ} (hΞ΄ : 0 < Ξ΄) : P {Ο‰ | βˆƒ a, βˆƒ t < n, pullCount A a t Ο‰ β‰  0 ∧ √(2 * pullCount A a t Ο‰ * Οƒ2 * Real.log (1 / Ξ΄)) ≀ sumRewards A R a t Ο‰ - pullCount A a t Ο‰ * (Ξ½ a)[id]} ≀ diff --git a/LeanMachineLearning/SequentialLearning/Algorithm.lean b/LeanMachineLearning/SequentialLearning/Algorithm.lean index 3185a206..4508c870 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithm.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithm.lean @@ -747,16 +747,22 @@ the algorithm only sees the past rounds. Since `Unit` carries a unique probabili observation kernels of such an environment are all equal to `Kernel.const _ (Measure.dirac ())`, and the observation process of an algorithm-environment sequence is `noObs`. -/ +/-- Every probability measure on `Unit` is `Measure.dirac ()`. -/ +lemma Measure.eq_dirac_unit (ΞΌ : Measure Unit) [IsProbabilityMeasure ΞΌ] : + ΞΌ = Measure.dirac () := by + ext s hs + rcases Set.eq_empty_or_nonempty s with rfl | ⟨u, hu⟩ + Β· simp + Β· have hs_univ : s = Set.univ := Set.eq_univ_of_forall fun x ↦ by rwa [Subsingleton.elim x u] + simp [hs_univ] + /-- Every Markov kernel with codomain `Unit` is the constant kernel at `Measure.dirac ()`. -/ lemma Kernel.eq_const_dirac_unit {Ξ± : Type*} {mΞ± : MeasurableSpace Ξ±} (ΞΊ : Kernel Ξ± Unit) [IsMarkovKernel ΞΊ] : ΞΊ = Kernel.const Ξ± (Measure.dirac ()) := by - ext a s hs + ext a : 1 rw [Kernel.const_apply] - rcases Set.eq_empty_or_nonempty s with rfl | ⟨u, hu⟩ - Β· simp - Β· have hs_univ : s = Set.univ := Set.eq_univ_of_forall fun x ↦ by rwa [Subsingleton.elim x u] - simp [hs_univ] + exact Measure.eq_dirac_unit (ΞΊ a) /-- A random variable with values in `Unit` admits any Markov kernel as conditional distribution. -/ diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean b/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean index 3f377515..9c48b76c 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean @@ -97,12 +97,12 @@ variable [NeZero K] {Ξ½ : Kernel (Fin K) 𝓨} [IsMarkovKernel Ξ½] /-- The action chosen at time `n` is the action `n % K`. -/ lemma action_ae_eq (n : β„•) - (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (stationaryEnv Ξ½) P (n + 1)) : + (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P (n + 1)) : A n =ᡐ[P] fun _ ↦ nextAction K n := h.action_detAlgorithm_ae_eq n.lt_succ_self lemma action_zero - (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (stationaryEnv Ξ½) P 1) : + (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P 1) : A 0 =ᡐ[P] fun _ ↦ 0 := by filter_upwards [action_ae_eq 0 h] with Ο‰ hΟ‰ rw [hΟ‰] @@ -110,7 +110,7 @@ lemma action_zero /-- At time `K * m`, the number of times each action is chosen is equal to `m`. -/ lemma pullCount_mul (m : β„•) - (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (stationaryEnv Ξ½) P (K * m)) + (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P (K * m)) (a : Fin K) : pullCount A a (K * m) =ᡐ[P] fun _ ↦ m := by rw [Filter.EventuallyEq] @@ -126,14 +126,14 @@ lemma pullCount_mul (m : β„•) _ = m := sum_mod_range_mul (Nat.pos_of_neZero K) m a lemma pullCount_eq_one - (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (stationaryEnv Ξ½) P K) (a : Fin K) : + (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P K) (a : Fin K) : pullCount A a K =ᡐ[P] fun _ ↦ 1 := by suffices pullCount A a (K * 1) =ᡐ[P] fun _ ↦ 1 by simpa using this refine pullCount_mul 1 (P := P) (Ξ½ := Ξ½) (O := O) (Y := Y) ?_ a simpa lemma time_gt_of_pullCount_gt_one - (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (stationaryEnv Ξ½) P K) (a : Fin K) : + (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P K) (a : Fin K) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, 1 < pullCount A a n Ο‰ β†’ K < n := by filter_upwards [pullCount_eq_one h a] with h h_eq n hn rw [← h_eq] at hn @@ -141,7 +141,7 @@ lemma time_gt_of_pullCount_gt_one exact hn.not_ge (pullCount_mono _ h_lt _) lemma pullCount_pos_of_time_ge - (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (stationaryEnv Ξ½) P K) : + (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P K) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, K ≀ n β†’ βˆ€ b : Fin K, 0 < pullCount A b n Ο‰ := by have h_ae a := pullCount_eq_one h a simp_rw [Filter.EventuallyEq, ← ae_all_iff] at h_ae @@ -151,7 +151,7 @@ lemma pullCount_pos_of_time_ge exact pullCount_mono _ hn _ lemma pullCount_pos_of_pullCount_gt_one - (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (stationaryEnv Ξ½) P K) (a : Fin K) : + (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P K) (a : Fin K) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, 1 < pullCount A a n Ο‰ β†’ βˆ€ b : Fin K, 0 < pullCount A b n Ο‰ := by filter_upwards [time_gt_of_pullCount_gt_one h a, pullCount_pos_of_time_ge h] with Ο‰ h1 h2 n h_gt a exact h2 n (h1 n h_gt).le a diff --git a/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean b/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean index dd29a648..673b0f21 100644 --- a/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean +++ b/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean @@ -14,7 +14,7 @@ public import LeanMachineLearning.SequentialLearning.StationaryEnv A Bayesian stationary environment is an environment that draws a parameter `e : 𝓔` from a prior `Q` before the first round and then behaves like the stationary environment -`stationaryEnv (ΞΊ.sectR e)`. Concretely, `bayesStationaryEnv Q ΞΊ : Environment 𝓔 𝓐 𝓨` announces +`Environment.bandit (ΞΊ.sectR e)`. Concretely, `bayesStationaryEnv Q ΞΊ : Environment 𝓔 𝓐 𝓨` announces an observation `e` at every round and runs against `alg.comapObs (fun _ ↦ ())`, for an `alg : Algorithm Unit 𝓐 𝓨`, an algorithm that does not use the observation. @@ -27,7 +27,7 @@ an `alg : Algorithm Unit 𝓐 𝓨`, an algorithm that does not use the observat the sequences of actions `A : β„• β†’ Ξ© β†’ 𝓐` and feedbacks `Y : β„• β†’ Ξ© β†’ 𝓨` are generated by the algorithm `alg : Algorithm Unit 𝓐 𝓨` interacting with `bayesStationaryEnv Q ΞΊ`, which it sees through `Algorithm.comapObs (fun _ ↦ ())`. Equivalently, `A` and `Y` are generated by `alg` - interacting with the stationary environment `stationaryEnv (ΞΊ.sectR (E Ο‰))`. + interacting with the stationary environment `Environment.bandit (ΞΊ.sectR (E Ο‰))`. * `bayesTrajMeasure Q ΞΊ alg`: for any choice of probability measure `Q : Measure 𝓔`, Markov kernel `ΞΊ : Kernel (𝓔 Γ— 𝓐) 𝓨`, and algorithm `alg : Algorithm Unit 𝓐 𝓨`, provides a probability measure `P : Measure (β„• β†’ Round 𝓔 𝓐 𝓨)` on a space that carries `E`, `A`, and `Y` such that @@ -43,12 +43,13 @@ an `alg : Algorithm Unit 𝓐 𝓨`, an algorithm that does not use the observat * `IsAlgEnvSeq.isBayesAlgEnvSeq`: a run of `alg.comapObs (fun _ ↦ ())` against `bayesStationaryEnv Q ΞΊ` is a Bayesian algorithm-environment sequence for the announced parameter. * `ae_IsAlgEnvSeq h`: if `h : IsBayesAlgEnvSeq Q ΞΊ alg E A Y P`, for `Q`-almost every `e : 𝓔`, - `IsAlgEnvSeq O' A' Y' alg (stationaryEnv (ΞΊ.sectR e)) (condDistrib (trajectory _ A Y) E P e)` for - some sequence of actions `A' : β„• β†’ (β„• β†’ Round Unit 𝓐 𝓨) β†’ 𝓐` and sequence of feedbacks + `IsAlgEnvSeq O' A' Y' alg (Environment.bandit (ΞΊ.sectR e)) (condDistrib (trajectory _ A Y) E P e)` + for some sequence of actions `A' : β„• β†’ (β„• β†’ Round Unit 𝓐 𝓨) β†’ 𝓐` and sequence of feedbacks `Y' : β„• β†’ (β„• β†’ Round Unit 𝓐 𝓨) β†’ 𝓨`. Intuitively, if the observable trajectory is generated by an underlying parameter `e : 𝓔`, the measure that carries the `IsBayesAlgEnvSeq` structure reveals a - measure that carries an `IsAlgEnvSeq` structure under the environment `stationaryEnv (ΞΊ.sectR e)` - and the same algorithm. This allows transferring results from the `IsAlgEnvSeq` structure to the + measure that carries an `IsAlgEnvSeq` structure under the environment + `Environment.bandit (ΞΊ.sectR e)` and the same algorithm. This allows transferring results from + the `IsAlgEnvSeq` structure to the `IsBayesAlgEnvSeq` structure. -/ @@ -114,7 +115,7 @@ lemma measurable_announceHist (n : β„•) : that the parameter `E : Ξ© β†’ 𝓔` has law `Q` and that the sequences of actions `A : β„• β†’ Ξ© β†’ 𝓐` and feedbacks `Y : β„• β†’ Ξ© β†’ 𝓨` are generated by the algorithm `alg : Algorithm Unit 𝓐 𝓨` interacting with an underlying environment that depends on `E` and `ΞΊ` -(`stationaryEnv (ΞΊ.sectR (E Ο‰))`). +(`Environment.bandit (ΞΊ.sectR (E Ο‰))`). This is `IsAlgEnvSeq` for the announcing environment `bayesStationaryEnv Q ΞΊ` and the algorithm `alg.comapObs (fun _ ↦ ())` that ignores the announced parameter: the observation at every round @@ -254,7 +255,7 @@ lemma hasLaw_IT_hist (h : IsBayesAlgEnvSeq Q ΞΊ alg E A Y P) (n : β„•) : rw [← Kernel.map_apply _ (IT.measurable_hist n), he]⟩ lemma ae_IsAlgEnvSeq (h : IsBayesAlgEnvSeq Q ΞΊ alg E A Y P) : - βˆ€α΅ e βˆ‚Q, IsAlgEnvSeq IT.obs IT.action IT.feedback alg (stationaryEnv (ΞΊ.sectR e)) + βˆ€α΅ e βˆ‚Q, IsAlgEnvSeq IT.obs IT.action IT.feedback alg (Environment.bandit (ΞΊ.sectR e)) (condDistrib (trajectory (noObs Ξ©) A Y) E P e) := by filter_upwards [ae_all_iff.2 (hasCondDistrib_IT_obs h), ae_all_iff.2 (hasCondDistrib_IT_action h), diff --git a/LeanMachineLearning/SequentialLearning/DivergenceDecomposition.lean b/LeanMachineLearning/SequentialLearning/DivergenceDecomposition.lean index 3cad44da..d80f6b19 100644 --- a/LeanMachineLearning/SequentialLearning/DivergenceDecomposition.lean +++ b/LeanMachineLearning/SequentialLearning/DivergenceDecomposition.lean @@ -158,22 +158,22 @@ variable {O : β„• β†’ Ξ© β†’ Unit} {O' : β„• β†’ Ξ©' β†’ Unit} {alg : Algorithm {ΞΊ ΞΊ' : Kernel 𝓐 𝓨} [IsMarkovKernel ΞΊ] [IsMarkovKernel ΞΊ'] /-- Chain rule for histories of a single algorithm versus two stationary environments. -/ -lemma IsAlgEnvSeq.klDiv_map_history_compProd (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΊ) P) - (h' : IsAlgEnvSeq O' A' Y' alg (stationaryEnv ΞΊ') P') (M : β„•) : +lemma IsAlgEnvSeq.klDiv_map_history_compProd (h : IsAlgEnvSeq O A Y alg (Environment.bandit ΞΊ) P) + (h' : IsAlgEnvSeq O' A' Y' alg (Environment.bandit ΞΊ') P') (M : β„•) : klDiv (P.map (history O A Y M)) (P'.map (history O' A' Y' M)) = βˆ‘ t ∈ range M, klDiv (P.map (A t) βŠ—β‚˜ ΞΊ) (P.map (A t) βŠ—β‚˜ ΞΊ') := by rw [h.klDiv_map_history_stepKernel h'] refine sum_congr rfl fun t _ ↦ ?_ have h_obs := (h.hasCondDistrib_obs t).map_eq - rw [obs_stationaryEnv] at h_obs - rw [stepKernel_stationaryEnv, stepKernel_stationaryEnv, + rw [obs_bandit] at h_obs + rw [stepKernel_bandit, stepKernel_bandit, klDiv_compProd_compProd_compProd_prodMkLeft_eq_klDiv_comp_compProd, ← h_obs, ← (h.hasCondDistrib_action t).hasLaw_comp.map_eq] /-- Chain rule for histories of a single algorithm versus two stationary environments. -/ lemma IsAlgEnvSeq.klDiv_map_history [MeasurableSpace.CountablyGenerated 𝓨] - (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΊ) P) - (h' : IsAlgEnvSeq O' A' Y' alg (stationaryEnv ΞΊ') P') (M : β„•) : + (h : IsAlgEnvSeq O A Y alg (Environment.bandit ΞΊ) P) + (h' : IsAlgEnvSeq O' A' Y' alg (Environment.bandit ΞΊ') P') (M : β„•) : klDiv (P.map (history O A Y M)) (P'.map (history O' A' Y' M)) = βˆ‘ t ∈ range M, ∫⁻ Ο‰, klDiv (ΞΊ (A t Ο‰)) (ΞΊ' (A t Ο‰)) βˆ‚P := by rw [h.klDiv_map_history_compProd h'] @@ -182,8 +182,8 @@ lemma IsAlgEnvSeq.klDiv_map_history [MeasurableSpace.CountablyGenerated 𝓨] lintegral_map (measurable_klDiv_kernel ΞΊ ΞΊ') (h.measurable_action t)] /-- Chain rule for trajectories of a single algorithm versus two stationary environments. -/ -lemma IsAlgEnvSeq.klDiv_map_trajectory_compProd (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΊ) P) - (h' : IsAlgEnvSeq O' A' Y' alg (stationaryEnv ΞΊ') P') : +lemma IsAlgEnvSeq.klDiv_map_trajectory_compProd (h : IsAlgEnvSeq O A Y alg (Environment.bandit ΞΊ) P) + (h' : IsAlgEnvSeq O' A' Y' alg (Environment.bandit ΞΊ') P') : klDiv (P.map (trajectory O A Y)) (P'.map (trajectory O' A' Y')) = βˆ‘' t : β„•, klDiv (P.map (A t) βŠ—β‚˜ ΞΊ) (P.map (A t) βŠ—β‚˜ ΞΊ') := by rw [klDiv_map_trajectory_eq_iSup h.measurable_obs h.measurable_action h.measurable_feedback @@ -192,8 +192,8 @@ lemma IsAlgEnvSeq.klDiv_map_trajectory_compProd (h : IsAlgEnvSeq O A Y alg (stat /-- Chain rule for trajectories of a single algorithm versus two stationary environments. -/ lemma IsAlgEnvSeq.klDiv_map_trajectory [MeasurableSpace.CountablyGenerated 𝓨] - (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΊ) P) - (h' : IsAlgEnvSeq O' A' Y' alg (stationaryEnv ΞΊ') P') : + (h : IsAlgEnvSeq O A Y alg (Environment.bandit ΞΊ) P) + (h' : IsAlgEnvSeq O' A' Y' alg (Environment.bandit ΞΊ') P') : klDiv (P.map (trajectory O A Y)) (P'.map (trajectory O' A' Y')) = βˆ‘' t : β„•, ∫⁻ Ο‰, klDiv (ΞΊ (A t Ο‰)) (ΞΊ' (A t Ο‰)) βˆ‚P := by rw [h.klDiv_map_trajectory_compProd h'] diff --git a/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean b/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean index fdae6c06..34477026 100644 --- a/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean +++ b/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean @@ -47,24 +47,25 @@ variable {𝓐 𝓨 : Type*} {m𝓐 : MeasurableSpace 𝓐} {m𝓨 : MeasurableS /-- The evaluation environment where the feedback is given by evaluating a fixed measurable function `f` at the chosen action. -/ noncomputable def onlineEvalEnv (g : β„• β†’ 𝓐 β†’ 𝓨) (hg : βˆ€ n, Measurable (g n)) := - obliviousEnv (fun n ↦ Kernel.deterministic (g n) (hg n)) + Environment.banditSeq (fun n ↦ Kernel.deterministic (g n) (hg n)) instance : IsObliviousEnv (onlineEvalEnv g hg) := - ⟨⟨fun n ↦ Kernel.deterministic (g n) (hg n), fun _ ↦ inferInstance, fun _ ↦ rfl⟩⟩ + inferInstanceAs (IsObliviousEnv (Environment.banditSeq fun n ↦ Kernel.deterministic (g n) (hg n))) instance : IsDeterministicEnv (onlineEvalEnv g hg) where exists_f n := ⟨fun p ↦ g n p.2, by fun_prop, rfl⟩ @[simp] -lemma feedbackCondAction_onlineEvalEnv (n : β„•) : - feedbackCondAction (onlineEvalEnv g hg) n = Kernel.deterministic (g n) (hg n) := by +lemma feedbackCondObsAction_onlineEvalEnv (n : β„•) : + (onlineEvalEnv g hg).feedbackCondObsAction n + = Kernel.deterministic (fun p ↦ g n p.2) (by fun_prop) := by simp [onlineEvalEnv] @[simp] lemma feedbackFun_onlineEvalEnv [MeasurableSpace.SeparatesPoints 𝓨] (n : β„•) : feedbackFun (onlineEvalEnv g hg) n = fun p ↦ g n p.2 := by have h_eq := feedback_eq_deterministic (onlineEvalEnv g hg) n - simpa only [onlineEvalEnv, feedback_obliviousEnv, Kernel.prodMkLeft_deterministic, + simpa only [onlineEvalEnv, feedback_banditSeq, Kernel.prodMkLeft_deterministic, Kernel.deterministic_inj] using h_eq.symm @[simp] @@ -80,17 +81,17 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm Unit 𝓐 𝓨 {P : Measure Ξ©} [IsProbabilityMeasure P] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} -lemma hascondDistrib_feedback_onlineEvalEnv +lemma hasCondDistrib_feedback_onlineEvalEnv (h : IsAlgEnvSeq O A Y alg (onlineEvalEnv g hg) P) (n : β„•) : - HasCondDistrib (Y n) (A n) (Kernel.deterministic (g n) (hg n)) P := by - simpa using IsObliviousEnv.hasCondDistrib_feedback h n + HasCondDistrib (Y n) (A n) (Kernel.deterministic (g n) (hg n)) P := + h.hasCondDistrib_feedback_banditSeq n lemma feedback_onlineEvalEnv_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] (h : IsAlgEnvSeq O A Y alg (onlineEvalEnv g hg) P) (n : β„•) : Y n =ᡐ[P] g n ∘ A n := ae_eq_of_condDistrib_eq_deterministic (hg n) (h.measurable_action n).aemeasurable (h.measurable_feedback n).aemeasurable - (hascondDistrib_feedback_onlineEvalEnv h n).condDistrib_eq + (hasCondDistrib_feedback_onlineEvalEnv h n).condDistrib_eq lemma forall_feedback_onlineEvalEnv_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] (h : IsAlgEnvSeq O A Y alg (onlineEvalEnv g hg) P) : @@ -110,8 +111,10 @@ instance : IsObliviousEnv (evalEnv f hf) := by unfold evalEnv; infer_instance instance : IsDeterministicEnv (evalEnv f hf) := by unfold evalEnv; infer_instance @[simp] -lemma feedbackCondAction_evalEnv (n : β„•) : - feedbackCondAction (evalEnv f hf) n = Kernel.deterministic f hf := by simp [evalEnv] +lemma feedbackCondObsAction_evalEnv (n : β„•) : + (evalEnv f hf).feedbackCondObsAction n + = Kernel.deterministic (fun p ↦ f p.2) (by fun_prop) := by + simp [evalEnv] @[simp] lemma feedbackFunZero_evalEnv [MeasurableSpace.SeparatesPoints 𝓨] : @@ -128,9 +131,9 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm Unit 𝓐 𝓨 {P : Measure Ξ©} [IsProbabilityMeasure P] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} -lemma hascondDistrib_feedback_evalEnv (h : IsAlgEnvSeq O A Y alg (evalEnv f hf) P) (n : β„•) : - HasCondDistrib (Y n) (A n) (Kernel.deterministic f hf) P := by - simpa using IsObliviousEnv.hasCondDistrib_feedback h n +lemma hasCondDistrib_feedback_evalEnv (h : IsAlgEnvSeq O A Y alg (evalEnv f hf) P) (n : β„•) : + HasCondDistrib (Y n) (A n) (Kernel.deterministic f hf) P := + h.hasCondDistrib_feedback_banditSeq n lemma feedback_evalEnv_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] (h : IsAlgEnvSeq O A Y alg (evalEnv f hf) P) (n : β„•) : diff --git a/LeanMachineLearning/SequentialLearning/Means.lean b/LeanMachineLearning/SequentialLearning/Means.lean index 6207a34f..59e0c2d1 100644 --- a/LeanMachineLearning/SequentialLearning/Means.lean +++ b/LeanMachineLearning/SequentialLearning/Means.lean @@ -71,16 +71,25 @@ lemma means_zero (env : Environment π“ž 𝓐 𝓨) (O : β„• β†’ Ξ© β†’ π“ž) (A @[simp] lemma means_of_isObliviousEnv [IsObliviousEnv env] (O : β„• β†’ Ξ© β†’ π“ž) (A : β„• β†’ Ξ© β†’ 𝓐) (Y : β„• β†’ Ξ© β†’ 𝓨) (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : - env.means O A Y k n Ο‰ = (feedbackCondAction env n k)[id] := by - simp [Environment.means, Environment.measure, feedback_eq_feedbackCondAction] + env.means O A Y k n Ο‰ = (env.feedbackCondObsAction n (O n Ο‰, k))[id] := by + simp [Environment.means, Environment.measure, env.feedback_eq_comap_feedbackCondObsAction, + Kernel.comap_apply] -lemma means_obliviousEnv (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] +lemma means_obliviousEnv (ΞΌ : β„• β†’ Measure π“ž) [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] + (Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : + (obliviousEnv ΞΌ Ξ½).means O A Y k n Ο‰ = (Ξ½ n (O n Ο‰, k))[id] := by simp + +lemma means_stationaryEnv (ΞΌ : Measure π“ž) [IsProbabilityMeasure ΞΌ] (Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨) + [IsMarkovKernel Ξ½] (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : + (stationaryEnv ΞΌ Ξ½).means O A Y k n Ο‰ = (Ξ½ (O n Ο‰, k))[id] := by simp + +lemma means_banditSeq (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : - (obliviousEnv Ξ½).means O A Y k n Ο‰ = (Ξ½ n k)[id] := by simp + (Environment.banditSeq Ξ½).means O A Y k n Ο‰ = (Ξ½ n k)[id] := by simp -lemma means_stationaryEnv (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] +lemma means_bandit (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : - (stationaryEnv Ξ½).means O A Y k n Ο‰ = (Ξ½ k)[id] := by simp + (Environment.bandit Ξ½).means O A Y k n Ο‰ = (Ξ½ k)[id] := by simp @[fun_prop] lemma IsAlgEnvSeq.stronglyMeasurable_means [SecondCountableTopology 𝓨] [OpensMeasurableSpace 𝓨] diff --git a/LeanMachineLearning/SequentialLearning/README.md b/LeanMachineLearning/SequentialLearning/README.md index a2fa4af1..d7a0bd64 100644 --- a/LeanMachineLearning/SequentialLearning/README.md +++ b/LeanMachineLearning/SequentialLearning/README.md @@ -28,11 +28,8 @@ In many applications, some of those kernels are deterministic, or do not depend We detail here the naming conventions for the various constructors, predicates and accessors that are used in the library. Generic constructions live in the `Algorithm` and `Environment` namespaces. -A `det` prefix marks the deterministic version of a constructor. Predicates are root-level `Is…Alg` / `Is…Env` classes when they carry an accessor, and namespaced `Prop` definitions otherwise. Accessors are namespaced so that dot notation works. -Every row may read the time `n` unless it says "stationary" or "nothing". -Every `det…` constructor is the stochastic one applied to `Kernel.deterministic`, gets both the determinism instance and the dependence instance, and takes its measurability proofs as `by fun_prop` autoparams. All time zero accessors are root-level `…0` definitions. Example: `Algorithm.policy0`. @@ -78,7 +75,7 @@ It can depend on time, history and action and be stochastic or deterministic. general: use `Environment` with Obs = Unit. -No history: use oblivious or stationary environment with Obs = Unit. +No history: `Environment.banditSeq` and `Environment.bandit` (no time). No history, deterministic: `Environment.evalSeq` and `Environment.eval` (no time). @@ -88,70 +85,6 @@ Nothing: `Environment.const` (stochastic). The deterministic version is probably ## Examples -Oblivious adversarial bandit environment: Obs = Unit, feedback depends on time and action, deterministic. Which constructor? `Environment.evalSeq`. +Oblivious adversarial bandit environment: Obs = Unit, feedback depends on time and action, deterministic. Use `Environment.evalSeq`. -Stochastic optimization: Obs = Unit, feedback depends on time and action, stochastic. Which constructor? `Environment.banditSeq`. - -| Policy reads | Stochastic constructor | Deterministic constructor | Predicate | Accessor | -|---|---|---|---|---| -| history, obs | `Algorithm` | `Algorithm.deterministic f`, was `detAlgorithm` | `IsDeterministicAlg`, exists | `alg.nextAction n`, was `nextAction alg n` | -| history only | `alg.comapObs fun _ ↦ ()`, exists | same | `Algorithm.IgnoresObs`, new def | none | -| obs | `Algorithm.markov Ο€`, `Ο€ : β„• β†’ Kernel π“ž 𝓐`, new | `Algorithm.detMarkov f`, `f : β„• β†’ π“ž β†’ 𝓐`, new | `IsMarkovAlg`, new | `alg.policyCondObs n : Kernel π“ž 𝓐` | -| obs, stationary | `Algorithm.stationary Ο€`, `Ο€ : Kernel π“ž 𝓐`, new | `Algorithm.detStationary f`, `f : π“ž β†’ 𝓐`, new | `IsStationaryAlg`, new | same, constant in `n` | -| time only | `Algorithm.openLoop ΞΌ`, `ΞΌ : β„• β†’ Measure 𝓐`, new | `Algorithm.ofSeq x`, was `fixedDesignAlg` | `IsOpenLoopAlg`, new | `alg.actionLaw n : Measure 𝓐` | -| nothing | `Algorithm.const ΞΌ`, was `randomSampling` | `Algorithm.detConst a`, new | none, use `IsOpenLoopAlg` | none | - -Named instances of rows: `Algorithm.uniform`, was `uniformAlgorithm`, is `const` of the uniform -measure; `Algorithm.roundRobin hK`, was `roundRobinAlgorithm`, is `ofSeq`. The bandit-specific -`ucbAlgorithm`, `etcAlgorithm` and `tsAlgorithm` keep their names in the `Bandits` namespace. - -Implications provided as instances: stationary implies Markov, open loop implies Markov, and the -deterministic constructors of each row give both instances. The three dependence predicates are -the named cases of `Algorithm.FactorsThrough`. - -## Environments - -### General observation type - -| Obs reads | Feedback reads | Stochastic constructor | Deterministic constructor | Predicate | Accessors | -|---|---|---|---|---|---| -| history | history, obs, action | the structure | `Environment.deterministic g f`, new | `IsDeterministicEnv`, redefined to cover both kernels | `env.obsFun n`, `env.feedbackFun n` | -| history | history, obs | `Environment.adversary ρ ΞΊ`, new | `Environment.detAdversary g f`, new | `Environment.FeedbackIgnoresAction`, from LMLPapers | none | -| last round | obs, action | `Environment.markov ρ P Ξ½`, new | `Environment.detMarkov sβ‚€ P f`, new | `IsMarkovEnv`, new | `env.transition n : Kernel (Round π“ž 𝓐 𝓨) π“ž`, `env.feedbackCondObsAction n : Kernel (π“ž Γ— 𝓐) 𝓨` | -| last obs and action, stationary | obs, action | `Environment.mdp ρ P r`, new | `Environment.detMdp sβ‚€ P r`, new | none, use `IsMarkovEnv` | same | -| time only | obs, action | `Environment.oblivious ρ Ξ½`, `ρ : β„• β†’ Measure π“ž`, `Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨`, new signature | `Environment.detOblivious o f`, new | `IsObliviousEnv`, redefined | `env.obsLaw n : Measure π“ž`, `env.feedbackCondObsAction n` | -| nothing | obs, action, stationary | `Environment.stationary ρ Ξ½`, `ρ : Measure π“ž`, `Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨`, new signature | `Environment.detStationary oβ‚€ f`, new | `IsStationaryEnv`, new | same, constant in `n` | - -Today's `detEnvironment obs f`, with random observations and deterministic feedback, becomes -`Environment.detFeedback obs f` with the definition `Environment.HasDeterministicFeedback`, or is -dropped since nothing uses it. - -### No observation - -With `π“ž = Unit` only the feedback matters, so these constructors are named by the feedback's shape -rather than by a dependence word. They are the existing bandit constructors, unchanged in type. - -| Feedback reads | Stochastic constructor | Deterministic constructor | -|---|---|---| -| time, action | `Environment.banditSeq Ξ½`, `Ξ½ : β„• β†’ Kernel 𝓐 𝓨`, was `obliviousEnv` | `Environment.evalSeq g`, was `onlineEvalEnv` | -| action | `Environment.bandit Ξ½`, `Ξ½ : Kernel 𝓐 𝓨`, was `stationaryEnv` | `Environment.eval f`, was `evalEnv` | -| time only | `Environment.indep P`, `P : β„• β†’ Measure 𝓨`, new | `Environment.ofSeq y`, was `seqEnv` in LMLPapers | -| nothing | `Environment.iid P`, was `iidEnv` in LMLPapers | `Environment.const y`, new | - -The general predicates apply to this table as they are. The one extra accessor is -`env.feedbackCondAction n : Kernel 𝓐 𝓨`, defined only for `π“ž = Unit` from -`feedbackCondObsAction`; it is today's `feedbackCondAction env n`. The identities tying the two -tables together are simp lemmas: `bandit Ξ½` is `stationary` of the Dirac measure and -`Ξ½.prodMkLeft Unit`, `eval f` is `detStationary`, `ofSeq y` is `evalSeq` of constant functions, -`iid P` is `bandit` of the constant kernel, and `indep P` is `banditSeq` of constant kernels. - -### Hidden parameter - -`Environment.bayes Q env : Environment (𝓔 Γ— π“ž) 𝓐 𝓨` for a family `env : 𝓔 β†’ Environment π“ž 𝓐 𝓨`, -new; `Environment.bayesBandit Q ΞΊ : Environment 𝓔 𝓐 𝓨`, was `bayesStationaryEnv`. - -Implications provided as instances: stationary implies oblivious implies Markov, and `mdp` is a -`markov` instance. - -Time zero: `Environment.obs0` unchanged; `Environment.feedback0`, was `Ξ½0`; -`Environment.feedbackFun0`, was `feedbackFunZero`. +Stochastic optimization: Obs = Unit, feedback depends on time and action, stochastic. Use `Environment.banditSeq`. diff --git a/LeanMachineLearning/SequentialLearning/StationaryEnv.lean b/LeanMachineLearning/SequentialLearning/StationaryEnv.lean index b652c7e5..1bc11152 100644 --- a/LeanMachineLearning/SequentialLearning/StationaryEnv.lean +++ b/LeanMachineLearning/SequentialLearning/StationaryEnv.lean @@ -11,29 +11,36 @@ public import LeanMachineLearning.SequentialLearning.Algorithm /-! # Oblivious and stationary environments -An oblivious environment is an environment in which the distribution of the next feedback depends -only on the last action (and not on the past history nor on the current observation). -If the kernel that gives the distribution of the next feedback given the last action is the same at -every time step, then we say that the environment is stationary. +An oblivious environment is an environment in which the distributions of the observation and of +the feedback do not depend on the past history: at time `n`, the observation has law +`env.obsLaw n`, and the feedback depends only on the current observation and action, through the +Markov kernel `env.feedbackCondObsAction n`. +If there are no observations (`π“ž = Unit`) and the kernel that gives the distribution of the +feedback given the action is the same at every time step, then we say that the environment is +stationary. ## Main definitions We define a `Prop`-valued typeclass `IsObliviousEnv` to express that an environment is oblivious, -and we define two constructors for oblivious environments. Those constructors build environments -without observations, that is with observation type `Unit`. +and we define constructors for oblivious environments, with and without observations. Typeclass and related definitions: * `IsObliviousEnv env`: the environment `env` is oblivious. -* `feedbackCondAction env n`: the kernel representing the conditional distribution of the feedback - given the action at time `n` in an oblivious environment `env`. +* `Environment.obsLaw env n`: the law of the observation at time `n` in an oblivious + environment `env`. +* `Environment.feedbackCondObsAction env n`: the kernel representing the conditional distribution + of the feedback given the observation and the action at time `n` in an oblivious + environment `env`. Constructors for oblivious environments: -* `obliviousEnv Ξ½`: an oblivious environment without observations, in which the distribution of the - next feedback depends only on the last action, but in a possibly time-dependent manner, and is - given by a sequence of Markov kernels `Ξ½ : β„• β†’ Kernel 𝓐 𝓨`. -* `stationaryEnv Ξ½`: a stationary environment without observations, in which the distribution of - the next feedback depends only on the last action (and not on the past history), and is given by - a Markov kernel `Ξ½ : Kernel 𝓐 𝓨`. +* `obliviousEnv ΞΌ Ξ½`: the oblivious environment in which the observation at time `n` has law `ΞΌ n` + and the feedback at time `n` is drawn from the Markov kernel `Ξ½ n : Kernel (π“ž Γ— 𝓐) 𝓨` applied to + the observation and the action at time `n`. +* `stationaryEnv ΞΌ Ξ½`: the oblivious environment with constant sequences: the observations have + law `ΞΌ` and the feedback is drawn from `Ξ½` applied to the observation and the action. +* `Environment.banditSeq Ξ½`, `Environment.bandit Ξ½`: the versions without observations + (`π“ž = Unit`), in which the feedback at time `n` is drawn from `Ξ½ n : Kernel 𝓐 𝓨` + (respectively from `Ξ½ : Kernel 𝓐 𝓨`) applied to the action at time `n`. -/ @@ -48,246 +55,480 @@ namespace Learning variable {π“ž 𝓐 𝓨 : Type*} {mπ“ž : MeasurableSpace π“ž} {m𝓐 : MeasurableSpace 𝓐} {m𝓨 : MeasurableSpace 𝓨} -/-- An environment is oblivious if the distribution of the next feedback depends only on -the last action and not on the past history nor on the current observation. -/ +/-- An environment is oblivious if the distributions of the next observation and feedback +don't depend on the past history: the observation at time `n` has a fixed law, and the feedback +at time `n` depends only on the observation and the action at time `n`. -/ class IsObliviousEnv (env : Environment π“ž 𝓐 𝓨) : Prop where - exists_eq_prodMkLeft : βˆƒ Ξ½ : β„• β†’ Kernel 𝓐 𝓨, (βˆ€ n, IsMarkovKernel (Ξ½ n)) ∧ - (βˆ€ n, env.feedback n = (Ξ½ n).prodMkLeft _) + exists_obs_eq_const : βˆƒ ΞΌ : β„• β†’ Measure π“ž, (βˆ€ n, IsProbabilityMeasure (ΞΌ n)) ∧ + βˆ€ n, env.obs n = Kernel.const _ (ΞΌ n) + exists_feedback_eq_comap : βˆƒ Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨, (βˆ€ n, IsMarkovKernel (Ξ½ n)) ∧ + βˆ€ n, env.feedback n = (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) -/-- The kernel representing the conditional distribution of the feedback given the action -at time `n` in an oblivious environment. -/ +namespace Environment + +/-- The law of the observation at time `n` in an oblivious environment. -/ noncomputable -def feedbackCondAction (env : Environment π“ž 𝓐 𝓨) [h_obl : IsObliviousEnv env] (n : β„•) : - Kernel 𝓐 𝓨 := - h_obl.exists_eq_prodMkLeft.choose n +def obsLaw (env : Environment π“ž 𝓐 𝓨) [h_obl : IsObliviousEnv env] (n : β„•) : Measure π“ž := + h_obl.exists_obs_eq_const.choose n instance (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] (n : β„•) : - IsMarkovKernel (feedbackCondAction env n) := - IsObliviousEnv.exists_eq_prodMkLeft.choose_spec.1 n + IsProbabilityMeasure (env.obsLaw n) := + IsObliviousEnv.exists_obs_eq_const.choose_spec.1 n + +lemma obs_eq_const_obsLaw (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] (n : β„•) : + env.obs n = Kernel.const _ (env.obsLaw n) := + IsObliviousEnv.exists_obs_eq_const.choose_spec.2 n -lemma feedback_eq_feedbackCondAction (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] (n : β„•) : - env.feedback n = (feedbackCondAction env n).prodMkLeft _ := - IsObliviousEnv.exists_eq_prodMkLeft.choose_spec.2 n +lemma obs0_eq_obsLaw (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] : + env.obs0 = env.obsLaw 0 := by + rw [Environment.obs0_def, obs_eq_const_obsLaw, Kernel.const_apply] + +/-- The kernel representing the conditional distribution of the feedback given the observation and +the action at time `n` in an oblivious environment. -/ +noncomputable +def feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [h_obl : IsObliviousEnv env] (n : β„•) : + Kernel (π“ž Γ— 𝓐) 𝓨 := + h_obl.exists_feedback_eq_comap.choose n -lemma Ξ½0_eq_feedbackCondAction (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] : - env.Ξ½0 = (feedbackCondAction env 0).prodMkLeft π“ž := by +instance (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] (n : β„•) : + IsMarkovKernel (env.feedbackCondObsAction n) := + IsObliviousEnv.exists_feedback_eq_comap.choose_spec.1 n + +lemma feedback_eq_comap_feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] + (n : β„•) : + env.feedback n = (env.feedbackCondObsAction n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := + IsObliviousEnv.exists_feedback_eq_comap.choose_spec.2 n + +lemma Ξ½0_eq_feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] : + env.Ξ½0 = env.feedbackCondObsAction 0 := by ext p : 1 - rw [Environment.Ξ½0_def, Kernel.comap_apply, feedback_eq_feedbackCondAction, - Kernel.prodMkLeft_apply, Kernel.prodMkLeft_apply] + rw [Environment.Ξ½0_def, Kernel.comap_apply, feedback_eq_comap_feedbackCondObsAction, + Kernel.comap_apply] + +end Environment namespace IsObliviousEnv variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} - {alg : Algorithm π“ž 𝓐 𝓨} {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} [IsFiniteMeasure P] + {alg : Algorithm π“ž 𝓐 𝓨} {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} {n N : β„•} - {Ξ½ : β„• β†’ Kernel 𝓐 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] + +/-- The observation at time `n` has law `env.obsLaw n`. -/ +lemma hasLaw_obs [IsProbabilityMeasure P] [IsObliviousEnv env] + (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + HasLaw (O n) (env.obsLaw n) P := by + have h' := h.hasCondDistrib_obs n + rw [env.obs_eq_const_obsLaw] at h' + exact h'.hasLaw_of_const + +variable [IsFiniteMeasure P] lemma hasCondDistrib_feedback_history_action [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : HasCondDistrib (Y n) (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) - ((feedbackCondAction env n).prodMkLeft _) P := by - rw [← feedback_eq_feedbackCondAction] + ((env.feedbackCondObsAction n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) + : Kernel ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐) 𝓨) P := by + rw [← env.feedback_eq_comap_feedbackCondObsAction] exact h.hasCondDistrib_feedback n +/-- The conditional distribution of the feedback at time `n` given the observation and the action +at time `n` is `env.feedbackCondObsAction n`. -/ lemma hasCondDistrib_feedback [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : - HasCondDistrib (Y n) (A n) (feedbackCondAction env n) P := + HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) (env.feedbackCondObsAction n) P := (hasCondDistrib_feedback_history_action h n).comp_right -/-- Conditionally on an event determined by the history before time `n` and the action at time -`n`, on which that action is equal to `b`, the feedback at time `n` has law -`feedbackCondAction env n b`. -/ +/-- Conditionally on an event determined by the history before time `n`, the observation and the +action at time `n`, on which the observation-action pair is equal to `b`, the feedback at time `n` +has law `env.feedbackCondObsAction n b`. -/ lemma hasLaw_feedback_cond [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) - {s : Set ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐)} (hs : MeasurableSet s) {b : 𝓐} (hsb : βˆ€ u ∈ s, u.2 = b) + {s : Set ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐)} (hs : MeasurableSet s) {b : π“ž Γ— 𝓐} + (hsb : βˆ€ u ∈ s, (u.1.2, u.2) = b) (hP : P ((fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s) β‰  0) : - HasLaw (Y n) (feedbackCondAction env n b) + HasLaw (Y n) (env.feedbackCondObsAction n b) P[|(fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s] := by refine (hasCondDistrib_feedback_history_action h n).hasLaw_cond (h.measurable_feedback _) hs (fun u hu ↦ ?_) hP - rw [Kernel.prodMkLeft_apply, hsb u hu] + rw [Kernel.comap_apply, hsb u hu] -/-- Conditionally on an event determined by the history before time `n` and the action at time -`n`, on which that action is constant, the feedback at time `n` is independent of the -history before time `n` and of the action at time `n`. -/ +/-- Conditionally on an event determined by the history before time `n`, the observation and the +action at time `n`, on which the observation-action pair is constant, the feedback at time `n` is +independent of the history before time `n`, the observation and the action at time `n`. -/ lemma indepFun_history_action_feedback_cond [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) - {s : Set ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐)} (hs : MeasurableSet s) {b : 𝓐} (hsb : βˆ€ u ∈ s, u.2 = b) : + {s : Set ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐)} (hs : MeasurableSet s) {b : π“ž Γ— 𝓐} + (hsb : βˆ€ u ∈ s, (u.1.2, u.2) = b) : (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) βŸ‚α΅’[P[|(fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s]] Y n := by have hO := h.measurable_obs have hA := h.measurable_action have hY := h.measurable_feedback refine (hasCondDistrib_feedback_history_action h n).indepFun_cond (by fun_prop) hs - (Ξ· := feedbackCondAction env n b) fun u hu ↦ ?_ - rw [Kernel.prodMkLeft_apply, hsb u hu] + (Ξ· := env.feedbackCondObsAction n b) fun u hu ↦ ?_ + rw [Kernel.comap_apply, hsb u hu] variable [StandardBorelSpace π“ž] [Nonempty π“ž] [StandardBorelSpace 𝓐] [Nonempty 𝓐] [StandardBorelSpace 𝓨] [Nonempty 𝓨] -/-- The feedback at time `n` is conditionally independent of the history before time `n` and of -the observation at time `n`, given the action at time `n`. -/ -lemma condIndepFun_feedback_history_action [StandardBorelSpace Ξ©] - [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : - Y n βŸ‚α΅’[A n, h.measurable_action _ ; P] (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) := by +/-- The feedback at time `n` is conditionally independent of the history before time `n`, given +the observation and the action at time `n`. -/ +lemma condIndepFun_feedback_history [StandardBorelSpace Ξ©] + [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + Y n βŸ‚α΅’[fun Ο‰ ↦ (O n Ο‰, A n Ο‰), (h.measurable_obs n).prodMk (h.measurable_action n); P] + history O A Y n := by have hO := h.measurable_obs have hA := h.measurable_action have hY := h.measurable_feedback refine condIndepFun_of_exists_condDistrib_prod_ae_eq_prodMkLeft - (Ξ· := feedbackCondAction env n) - (by fun_prop) (by fun_prop) (by fun_prop) ?_ + (Ξ· := env.feedbackCondObsAction n) (by fun_prop) (by fun_prop) (by fun_prop) ?_ refine HasCondDistrib.condDistrib_eq ?_ - rw [← feedback_eq_feedbackCondAction] - exact h.hasCondDistrib_feedback n - -lemma condIndepFun_feedback_history_action_action [StandardBorelSpace Ξ©] + have h' := hasCondDistrib_feedback_history_action h n + have hΞΊ : ((env.feedbackCondObsAction n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) + : Kernel ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐) 𝓨) + = ((env.feedbackCondObsAction n).prodMkLeft (Hist π“ž 𝓐 𝓨 n)).comap + (fun p ↦ (p.1.1, (p.1.2, p.2))) (by fun_prop) := rfl + rw [hΞΊ] at h' + exact h'.comp_right + +/-- The feedback at time `n` is conditionally independent of the history before time `n`, the +observation and the action at time `n`, given the observation and the action at time `n`. -/ +lemma condIndepFun_feedback_history_obs_action [StandardBorelSpace Ξ©] [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : - Y n βŸ‚α΅’[A n, h.measurable_action n; P] (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) := by - have h_indep : Y n βŸ‚α΅’[A n, h.measurable_action n; P] - (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) := condIndepFun_feedback_history_action h n + Y n βŸ‚α΅’[fun Ο‰ ↦ (O n Ο‰, A n Ο‰), (h.measurable_obs n).prodMk (h.measurable_action n); P] + (fun Ο‰ ↦ (history O A Y n Ο‰, (O n Ο‰, A n Ο‰))) := by have hO := h.measurable_obs have hA := h.measurable_action have hY := h.measurable_feedback - exact h_indep.prod_right (by fun_prop) (by fun_prop) (by fun_prop) + exact (condIndepFun_feedback_history h n).prod_right (by fun_prop) (by fun_prop) (by fun_prop) end IsObliviousEnv -/-- An oblivious environment without observations, in which the distribution of the next feedback -depends only on the last action, but in a possibly time-dependent manner. -/ -@[simps] +section Oblivious + +variable {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] + {Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨} [hΞ½ : βˆ€ n, IsMarkovKernel (Ξ½ n)] + +/-- The oblivious environment in which the observation at time `n` has law `ΞΌ n` and the feedback +at time `n` is drawn from `Ξ½ n` applied to the observation and the action at time `n`, whatever the +past history. -/ noncomputable -def obliviousEnv (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : Environment Unit 𝓐 𝓨 where - obs _ := Kernel.const _ (Measure.dirac ()) - feedback n := (Ξ½ n).prodMkLeft _ +def obliviousEnv (ΞΌ : β„• β†’ Measure π“ž) [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] + (Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : Environment π“ž 𝓐 𝓨 where + obs n := Kernel.const _ (ΞΌ n) + feedback n := (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) + +@[simp] +lemma obs_obliviousEnv (n : β„•) : (obliviousEnv ΞΌ Ξ½).obs n = Kernel.const _ (ΞΌ n) := rfl -lemma feedback_obliviousEnv (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] (n : β„•) : - (obliviousEnv Ξ½).feedback n = (Ξ½ n).prodMkLeft _ := rfl +@[simp] +lemma feedback_obliviousEnv (n : β„•) : + (obliviousEnv ΞΌ Ξ½).feedback n = (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := rfl @[simp] -lemma obs0_obliviousEnv (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : - (obliviousEnv Ξ½).obs0 = Measure.dirac () := rfl +lemma obs0_obliviousEnv : (obliviousEnv ΞΌ Ξ½).obs0 = ΞΌ 0 := rfl @[simp] -lemma Ξ½0_obliviousEnv (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : - (obliviousEnv Ξ½).Ξ½0 = (Ξ½ 0).prodMkLeft Unit := by +lemma Ξ½0_obliviousEnv : (obliviousEnv ΞΌ Ξ½).Ξ½0 = Ξ½ 0 := by ext p : 1 - rw [Environment.Ξ½0_def, Kernel.comap_apply, feedback_obliviousEnv, Kernel.prodMkLeft_apply, - Kernel.prodMkLeft_apply] + rw [Environment.Ξ½0_def, Kernel.comap_apply, feedback_obliviousEnv, Kernel.comap_apply] + +lemma stepKernel_obliviousEnv (alg : Algorithm π“ž 𝓐 𝓨) (n : β„•) : + stepKernel alg (obliviousEnv ΞΌ Ξ½) n + = Kernel.const _ (ΞΌ n) + βŠ—β‚– (alg.policy n βŠ—β‚– (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop)) := + rfl -instance (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : - IsObliviousEnv (obliviousEnv Ξ½) where - exists_eq_prodMkLeft := ⟨ν, inferInstance, fun _ ↦ rfl⟩ +instance : IsObliviousEnv (obliviousEnv ΞΌ Ξ½) where + exists_obs_eq_const := ⟨μ, inferInstance, fun _ ↦ rfl⟩ + exists_feedback_eq_comap := ⟨ν, inferInstance, fun _ ↦ rfl⟩ +/-- The law of the observations of `obliviousEnv ΞΌ Ξ½` is `ΞΌ`. The nonemptiness assumptions ensure +that there are histories of every length, so that the observation kernels determine `ΞΌ`. -/ @[simp] -lemma feedbackCondAction_obliviousEnv (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [hΞ½ : βˆ€ n, IsMarkovKernel (Ξ½ n)] - (n : β„•) : - feedbackCondAction (obliviousEnv Ξ½) n = Ξ½ n := by +lemma obsLaw_obliviousEnv [Nonempty 𝓐] [Nonempty 𝓨] (n : β„•) : + (obliviousEnv ΞΌ Ξ½).obsLaw n = ΞΌ n := by + have : Nonempty π“ž := Measure.nonempty_of_neZero (ΞΌ n) + have h_eq := (obliviousEnv ΞΌ Ξ½).obs_eq_const_obsLaw n + rw [obs_obliviousEnv, Kernel.ext_iff] at h_eq + simpa using (h_eq (Classical.arbitrary _)).symm + +@[simp] +lemma feedbackCondObsAction_obliviousEnv (n : β„•) : + (obliviousEnv ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ n := by + rcases isEmpty_or_nonempty π“ž with hπ“ž | hπ“ž + Β· ext p : 1 + exact hπ“ž.elim p.1 rcases isEmpty_or_nonempty 𝓐 with h𝓐 | h𝓐 - Β· ext a : 1 - exact h𝓐.elim a - rcases isEmpty_or_nonempty 𝓨 with hR | hR + Β· ext p : 1 + exact h𝓐.elim p.2 + rcases isEmpty_or_nonempty 𝓨 with h𝓨 | h𝓨 Β· refine absurd (hΞ½ 0) ?_ simp only [Subsingleton.eq_zero Ξ½, Pi.zero_apply] exact Kernel.not_isMarkovKernel_zero - have : Nonempty (Hist Unit 𝓐 𝓨 n Γ— Unit) := ⟨(fun _ ↦ ((), h𝓐.some, hR.some), ())⟩ - have h_eq := feedback_eq_feedbackCondAction (obliviousEnv Ξ½) n - rw [feedback_obliviousEnv, Kernel.prodMkLeft_inj] at h_eq - exact h_eq.symm + have h_eq := (obliviousEnv ΞΌ Ξ½).feedback_eq_comap_feedbackCondObsAction n + rw [feedback_obliviousEnv, Kernel.ext_iff] at h_eq + ext p : 1 + obtain ⟨o, a⟩ := p + exact (h_eq ((Classical.arbitrary _, o), a)).symm + +end Oblivious -/-- A stationary environment without observations, in which the distribution of the next feedback -depends only on the last action. -/ +section Stationary + +variable {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] {Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨} [IsMarkovKernel Ξ½] + +/-- The stationary environment in which the observations have law `ΞΌ` and the feedback is drawn +from `Ξ½` applied to the observation and the action, whatever the past history. -/ noncomputable -def stationaryEnv (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] : Environment Unit 𝓐 𝓨 := - obliviousEnv fun _ ↦ Ξ½ +def stationaryEnv (ΞΌ : Measure π“ž) [IsProbabilityMeasure ΞΌ] (Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨) + [IsMarkovKernel Ξ½] : Environment π“ž 𝓐 𝓨 := + obliviousEnv (fun _ ↦ ΞΌ) (fun _ ↦ Ξ½) + +lemma stationaryEnv_def : stationaryEnv ΞΌ Ξ½ = obliviousEnv (fun _ ↦ ΞΌ) (fun _ ↦ Ξ½) := rfl @[simp] -lemma obs_stationaryEnv (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] (n : β„•) : - (stationaryEnv Ξ½).obs n = Kernel.const _ (Measure.dirac ()) := rfl +lemma obs_stationaryEnv (n : β„•) : (stationaryEnv ΞΌ Ξ½).obs n = Kernel.const _ ΞΌ := rfl @[simp] -lemma feedback_stationaryEnv (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] (n : β„•) : - (stationaryEnv Ξ½).feedback n = Ξ½.prodMkLeft _ := rfl +lemma feedback_stationaryEnv (n : β„•) : + (stationaryEnv ΞΌ Ξ½).feedback n = Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := rfl -lemma stepKernel_stationaryEnv (alg : Algorithm Unit 𝓐 𝓨) (Ξ· : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ·] - (n : β„•) : - stepKernel alg (stationaryEnv Ξ·) n - = Kernel.const _ (Measure.dirac ()) βŠ—β‚– (alg.policy n βŠ—β‚– Ξ·.prodMkLeft _) := by - rw [stepKernel_def, obs_stationaryEnv, feedback_stationaryEnv] +@[simp] +lemma obs0_stationaryEnv : (stationaryEnv ΞΌ Ξ½).obs0 = ΞΌ := rfl + +@[simp] +lemma Ξ½0_stationaryEnv : (stationaryEnv ΞΌ Ξ½).Ξ½0 = Ξ½ := Ξ½0_obliviousEnv + +lemma stepKernel_stationaryEnv (alg : Algorithm π“ž 𝓐 𝓨) (n : β„•) : + stepKernel alg (stationaryEnv ΞΌ Ξ½) n + = Kernel.const _ ΞΌ βŠ—β‚– (alg.policy n βŠ—β‚– Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop)) := + rfl + +instance : IsObliviousEnv (stationaryEnv ΞΌ Ξ½) := + inferInstanceAs (IsObliviousEnv (obliviousEnv _ _)) + +@[simp] +lemma obsLaw_stationaryEnv [Nonempty 𝓐] [Nonempty 𝓨] (n : β„•) : + (stationaryEnv ΞΌ Ξ½).obsLaw n = ΞΌ := + obsLaw_obliviousEnv n + +@[simp] +lemma feedbackCondObsAction_stationaryEnv (n : β„•) : + (stationaryEnv ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ := + feedbackCondObsAction_obliviousEnv n + +end Stationary + +section BanditSeq + +variable {Ξ½ : β„• β†’ Kernel 𝓐 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] + +/-- The oblivious environment without observations in which the feedback at time `n` is drawn from +`Ξ½ n` applied to the action at time `n`, whatever the past history. -/ +noncomputable +def Environment.banditSeq (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : + Environment Unit 𝓐 𝓨 := + obliviousEnv (fun _ ↦ Measure.dirac ()) (fun n ↦ (Ξ½ n).prodMkLeft Unit) + +lemma Environment.banditSeq_def : + Environment.banditSeq Ξ½ + = obliviousEnv (fun _ ↦ Measure.dirac ()) (fun n ↦ (Ξ½ n).prodMkLeft Unit) := + rfl + +@[simp] +lemma obs_banditSeq (n : β„•) : + (Environment.banditSeq Ξ½).obs n = Kernel.const _ (Measure.dirac ()) := rfl + +@[simp] +lemma feedback_banditSeq (n : β„•) : (Environment.banditSeq Ξ½).feedback n = (Ξ½ n).prodMkLeft _ := rfl + +@[simp] +lemma obs0_banditSeq : (Environment.banditSeq Ξ½).obs0 = Measure.dirac () := rfl + +@[simp] +lemma Ξ½0_banditSeq : (Environment.banditSeq Ξ½).Ξ½0 = (Ξ½ 0).prodMkLeft Unit := Ξ½0_obliviousEnv + +instance : IsObliviousEnv (Environment.banditSeq Ξ½) := + inferInstanceAs (IsObliviousEnv (obliviousEnv _ _)) + +@[simp] +lemma obsLaw_banditSeq (n : β„•) : (Environment.banditSeq Ξ½).obsLaw n = Measure.dirac () := + Measure.eq_dirac_unit _ + +@[simp] +lemma feedbackCondObsAction_banditSeq (n : β„•) : + (Environment.banditSeq Ξ½).feedbackCondObsAction n = (Ξ½ n).prodMkLeft Unit := + feedbackCondObsAction_obliviousEnv n + +end BanditSeq + +section Bandit + +variable {Ξ½ : Kernel 𝓐 𝓨} [IsMarkovKernel Ξ½] + +/-- The stationary environment without observations in which the feedback is drawn from `Ξ½` +applied to the action, whatever the past history: a stochastic bandit. -/ +noncomputable +def Environment.bandit (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] : Environment Unit 𝓐 𝓨 := + stationaryEnv (Measure.dirac ()) (Ξ½.prodMkLeft Unit) + +lemma Environment.bandit_def : + Environment.bandit Ξ½ = stationaryEnv (Measure.dirac ()) (Ξ½.prodMkLeft Unit) := rfl + +lemma Environment.bandit_eq_banditSeq : Environment.bandit Ξ½ = Environment.banditSeq fun _ ↦ Ξ½ := + rfl + +@[simp] +lemma obs_bandit (n : β„•) : (Environment.bandit Ξ½).obs n = Kernel.const _ (Measure.dirac ()) := rfl + +@[simp] +lemma feedback_bandit (n : β„•) : (Environment.bandit Ξ½).feedback n = Ξ½.prodMkLeft _ := rfl + +lemma stepKernel_bandit (alg : Algorithm Unit 𝓐 𝓨) (n : β„•) : + stepKernel alg (Environment.bandit Ξ½) n + = Kernel.const _ (Measure.dirac ()) βŠ—β‚– (alg.policy n βŠ—β‚– Ξ½.prodMkLeft _) := by + rw [stepKernel_def, obs_bandit, feedback_bandit] @[simp] -lemma obs0_stationaryEnv (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] : - (stationaryEnv Ξ½).obs0 = Measure.dirac () := rfl +lemma obs0_bandit : (Environment.bandit Ξ½).obs0 = Measure.dirac () := rfl @[simp] -lemma Ξ½0_stationaryEnv (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] : - (stationaryEnv Ξ½).Ξ½0 = Ξ½.prodMkLeft Unit := - Ξ½0_obliviousEnv _ +lemma Ξ½0_bandit : (Environment.bandit Ξ½).Ξ½0 = Ξ½.prodMkLeft Unit := Ξ½0_obliviousEnv + +instance : IsObliviousEnv (Environment.bandit Ξ½) := + inferInstanceAs (IsObliviousEnv (obliviousEnv _ _)) -instance (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] : IsObliviousEnv (stationaryEnv Ξ½) where - exists_eq_prodMkLeft := ⟨fun _ ↦ Ξ½, inferInstance, fun _ ↦ rfl⟩ +@[simp] +lemma obsLaw_bandit (n : β„•) : (Environment.bandit Ξ½).obsLaw n = Measure.dirac () := + Measure.eq_dirac_unit _ @[simp] -lemma feedbackCondAction_stationaryEnv (Ξ½ : Kernel 𝓐 𝓨) [hΞ½ : IsMarkovKernel Ξ½] (n : β„•) : - feedbackCondAction (stationaryEnv Ξ½) n = Ξ½ := feedbackCondAction_obliviousEnv _ _ +lemma feedbackCondObsAction_bandit (n : β„•) : + (Environment.bandit Ξ½).feedbackCondObsAction n = Ξ½.prodMkLeft Unit := + feedbackCondObsAction_obliviousEnv n + +end Bandit + +namespace IsAlgEnvSeq + +section General + +variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm π“ž 𝓐 𝓨} + {P : Measure Ξ©} [IsProbabilityMeasure P] + {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} + +/-- The observation at time `n` has law `ΞΌ n`. -/ +lemma hasLaw_obs_obliviousEnv {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] + {Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] + (h : IsAlgEnvSeq O A Y alg (obliviousEnv ΞΌ Ξ½) P) (n : β„•) : + HasLaw (O n) (ΞΌ n) P := by + have h' := h.hasCondDistrib_obs n + rw [obs_obliviousEnv] at h' + exact h'.hasLaw_of_const + +/-- The conditional distribution of the feedback at time `n` given the observation and the action +at time `n` is `Ξ½ n`. -/ +lemma hasCondDistrib_feedback_obliviousEnv {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] + {Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] + (h : IsAlgEnvSeq O A Y alg (obliviousEnv ΞΌ Ξ½) P) (n : β„•) : + HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) (Ξ½ n) P := by + have h' := h.hasCondDistrib_feedback n + rw [feedback_obliviousEnv] at h' + exact h'.comp_right + +/-- The observation at time `n` has law `ΞΌ`. -/ +lemma hasLaw_obs_stationaryEnv {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] + {Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨} [IsMarkovKernel Ξ½] + (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΌ Ξ½) P) (n : β„•) : + HasLaw (O n) ΞΌ P := + hasLaw_obs_obliviousEnv h n + +/-- The conditional distribution of the feedback at time `n` given the observation and the action +at time `n` is `Ξ½`. -/ +lemma hasCondDistrib_feedback_stationaryEnv {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] + {Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨} [IsMarkovKernel Ξ½] + (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΌ Ξ½) P) (n : β„•) : + HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) Ξ½ P := + hasCondDistrib_feedback_obliviousEnv h n + +end General + +section Bandit variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm Unit 𝓐 𝓨} {Ξ½ : Kernel 𝓐 𝓨} [IsMarkovKernel Ξ½] {P : Measure Ξ©} [IsProbabilityMeasure P] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} -namespace IsAlgEnvSeq - /-- The conditional distribution of the feedback at time `n` given the action at time `n` is `Ξ½ n`. -/ -lemma hasCondDistrib_feedback_obliviousEnv {Ξ½ : β„• β†’ Kernel 𝓐 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] - (h : IsAlgEnvSeq O A Y alg (obliviousEnv Ξ½) P) (n : β„•) : +lemma hasCondDistrib_feedback_banditSeq {Ξ½ : β„• β†’ Kernel 𝓐 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] + (h : IsAlgEnvSeq O A Y alg (Environment.banditSeq Ξ½) P) (n : β„•) : HasCondDistrib (Y n) (A n) (Ξ½ n) P := by - simpa using IsObliviousEnv.hasCondDistrib_feedback h n + have h' : HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) ((Ξ½ n).prodMkLeft Unit) P := + hasCondDistrib_feedback_obliviousEnv h n + exact h'.comp_right /-- The conditional distribution of the feedback at time `n` given the action at time `n` is `Ξ½`. -/ -lemma hasCondDistrib_feedback_stationaryEnv - (h : IsAlgEnvSeq O A Y alg (stationaryEnv Ξ½) P) (n : β„•) : +lemma hasCondDistrib_feedback_bandit + (h : IsAlgEnvSeq O A Y alg (Environment.bandit Ξ½) P) (n : β„•) : HasCondDistrib (Y n) (A n) Ξ½ P := - hasCondDistrib_feedback_obliviousEnv h n + hasCondDistrib_feedback_banditSeq h n /-- The conditional distribution of the feedback at time `n` given the action at time `n` is `Ξ½`. -/ -lemma condDistrib_feedback_stationaryEnv [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (stationaryEnv Ξ½) P) (n : β„•) : +lemma condDistrib_feedback_bandit [StandardBorelSpace 𝓨] [Nonempty 𝓨] + (h : IsAlgEnvSeq O A Y alg (Environment.bandit Ξ½) P) (n : β„•) : condDistrib (Y n) (A n) P =ᡐ[P.map (A n)] Ξ½ := - (hasCondDistrib_feedback_stationaryEnv h n).condDistrib_eq + (hasCondDistrib_feedback_bandit h n).condDistrib_eq /-- Conditionally on an event determined by the history before time `n` and the action at time `n`, on which that action is equal to `b`, the feedback at time `n` has law `Ξ½ b`. -/ -lemma hasLaw_feedback_cond_stationaryEnv (h : IsAlgEnvSeq O A Y alg (stationaryEnv Ξ½) P) (n : β„•) +lemma hasLaw_feedback_cond_bandit (h : IsAlgEnvSeq O A Y alg (Environment.bandit Ξ½) P) (n : β„•) {s : Set ((Hist Unit 𝓐 𝓨 n Γ— Unit) Γ— 𝓐)} (hs : MeasurableSet s) {b : 𝓐} (hsb : βˆ€ u ∈ s, u.2 = b) (hP : P ((fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s) β‰  0) : HasLaw (Y n) (Ξ½ b) P[|(fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s] := by - simpa using IsObliviousEnv.hasLaw_feedback_cond h n hs hsb hP + simpa using IsObliviousEnv.hasLaw_feedback_cond h n hs (b := ((), b)) + (fun u hu ↦ by simp [hsb u hu]) hP /-- Conditionally on an event determined by the history before time `n` and the action at time `n`, on which that action is constant, the feedback at time `n` is independent of the history before time `n` and of the action at time `n`. -/ -lemma indepFun_history_action_feedback_cond_stationaryEnv - (h : IsAlgEnvSeq O A Y alg (stationaryEnv Ξ½) P) (n : β„•) +lemma indepFun_history_action_feedback_cond_bandit + (h : IsAlgEnvSeq O A Y alg (Environment.bandit Ξ½) P) (n : β„•) {s : Set ((Hist Unit 𝓐 𝓨 n Γ— Unit) Γ— 𝓐)} (hs : MeasurableSet s) {b : 𝓐} (hsb : βˆ€ u ∈ s, u.2 = b) : (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) βŸ‚α΅’[P[|(fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s]] Y n := - IsObliviousEnv.indepFun_history_action_feedback_cond h n hs hsb + IsObliviousEnv.indepFun_history_action_feedback_cond h n hs (b := ((), b)) + fun u hu ↦ by simp [hsb u hu] /-- The feedback at time `n` is conditionally independent of the history before time `n` given the action at time `n`. -/ -lemma condIndepFun_feedback_history_action [StandardBorelSpace Ξ©] +lemma condIndepFun_feedback_history_action_bandit [StandardBorelSpace Ξ©] [StandardBorelSpace 𝓐] [Nonempty 𝓐] [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (stationaryEnv Ξ½) P) (n : β„•) : - Y n βŸ‚α΅’[A n, h.measurable_action _ ; P] (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) := - IsObliviousEnv.condIndepFun_feedback_history_action h n + (h : IsAlgEnvSeq O A Y alg (Environment.bandit Ξ½) P) (n : β„•) : + Y n βŸ‚α΅’[A n, h.measurable_action _ ; P] (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) := by + have hO := h.measurable_obs + have hA := h.measurable_action + have hY := h.measurable_feedback + refine condIndepFun_of_exists_condDistrib_prod_ae_eq_prodMkLeft (Ξ· := Ξ½) + (by fun_prop) (by fun_prop) (by fun_prop) ?_ + refine HasCondDistrib.condDistrib_eq ?_ + have h' := h.hasCondDistrib_feedback n + rwa [feedback_bandit] at h' -lemma condIndepFun_feedback_history_action_action [StandardBorelSpace Ξ©] +lemma condIndepFun_feedback_history_action_action_bandit [StandardBorelSpace Ξ©] [StandardBorelSpace 𝓐] [Nonempty 𝓐] [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (stationaryEnv Ξ½) P) (n : β„•) : + (h : IsAlgEnvSeq O A Y alg (Environment.bandit Ξ½) P) (n : β„•) : Y n βŸ‚α΅’[A n, h.measurable_action n; P] - (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) := - IsObliviousEnv.condIndepFun_feedback_history_action_action h n + (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) := by + have hO := h.measurable_obs + have hA := h.measurable_action + have hY := h.measurable_feedback + exact (condIndepFun_feedback_history_action_bandit h n).prod_right (by fun_prop) (by fun_prop) + (by fun_prop) + +end Bandit end IsAlgEnvSeq From f6e852def5c25763568a52cdadc01f00d5a4eb37 Mon Sep 17 00:00:00 2001 From: Remy Degenne Date: Thu, 17 Sep 2026 14:06:54 +0200 Subject: [PATCH 3/7] add markov algorithm --- LeanMachineLearning.lean | 1 + .../SequentialLearning/Algorithms/Markov.lean | 258 ++++++++++++++++++ .../Algorithms/RandomSampling/Basic.lean | 18 +- .../Algorithms/Uniform.lean | 3 + 4 files changed, 277 insertions(+), 3 deletions(-) diff --git a/LeanMachineLearning.lean b/LeanMachineLearning.lean index 3d683931..f605d389 100644 --- a/LeanMachineLearning.lean +++ b/LeanMachineLearning.lean @@ -51,6 +51,7 @@ public import LeanMachineLearning.SequentialLearning.ActionIndicator public import LeanMachineLearning.SequentialLearning.Algorithm public import LeanMachineLearning.SequentialLearning.AlgorithmDensity public import LeanMachineLearning.SequentialLearning.AlgorithmDensityBayes +public import LeanMachineLearning.SequentialLearning.Algorithms.Markov public import LeanMachineLearning.SequentialLearning.Algorithms.RandomSampling.Basic public import LeanMachineLearning.SequentialLearning.Algorithms.RandomSampling.Tendsto public import LeanMachineLearning.SequentialLearning.Algorithms.RoundRobin diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean b/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean index e69de29b..f8872080 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean @@ -0,0 +1,258 @@ +/- +Copyright (c) 2025 RΓ©my Degenne. All rights reserved. +Released under Apache 2.0 license as described in the file LICENSE. +Authors: RΓ©my Degenne, Paulo Rauber +-/ +module + +public import LeanMachineLearning.ForMathlib.Probability.Kernel.Composition.MapComap +public import LeanMachineLearning.SequentialLearning.Algorithm + +/-! +# Algorithms that depend only on the observation + +In general, an algorithm `alg : Algorithm π“ž 𝓐 𝓨` has a policy +`alg.policy n : Kernel (Hist π“ž 𝓐 𝓨 n Γ— π“ž) 𝓐`, which depends on the history before time `n` and on +the observation at time `n`. In some cases, the policy depends only on the observation. +We say that `alg` is a *Markov algorithm* if there exists a Markov kernel `ΞΊ : Kernel π“ž 𝓐` such that +`alg.policy n = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n)` for all `n`. + +## Main definitions + +* `Algorithm.IsMarkov alg`: the policy of `alg` depends only on the current observation, + not on the history and not on the time. +* `Algorithm.policyCondObs alg`: the kernel representing the conditional distribution of the + action given the observation in a Markov algorithm `alg`. +* `Algorithm.markov ΞΊ`: an algorithm with a policy that depends only on the current observation, + given by the Markov kernel `ΞΊ`. + +## Main statements + +* `Algorithm.IsMarkov.hasCondDistrib_action`: in a run of a Markov algorithm, the conditional + distribution of the action at time `n` given the observation at time `n` is `alg.policyCondObs`. +* `Algorithm.IsMarkov.condIndepFun_action_history`: in a run of a Markov algorithm, the action at + time `n` is conditionally independent of the history before time `n` given the observation at + time `n`. + +-/ + +@[expose] public section + +open MeasureTheory ProbabilityTheory Filter Real Finset + +open scoped ENNReal NNReal + +namespace Learning + +variable {π“ž 𝓐 𝓨 : Type*} {mπ“ž : MeasurableSpace π“ž} {m𝓐 : MeasurableSpace 𝓐} + {m𝓨 : MeasurableSpace 𝓨} + +/-- The policy of the algorithm depends only on the current observation, not on the history and +not on the time. -/ +class Algorithm.IsMarkov (alg : Algorithm π“ž 𝓐 𝓨) : Prop where + exists_policy_eq_prodMkLeft : βˆƒ ΞΊ : Kernel π“ž 𝓐, βˆ€ n, alg.policy n = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n) + +namespace Algorithm + +/-- The kernel representing the conditional distribution of the action given the observation +in a Markov algorithm. -/ +noncomputable +def policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [h_markov : alg.IsMarkov] : Kernel π“ž 𝓐 := + h_markov.exists_policy_eq_prodMkLeft.choose + +lemma policy_eq_prodMkLeft_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] (n : β„•) : + alg.policy n = alg.policyCondObs.prodMkLeft (Hist π“ž 𝓐 𝓨 n) := + IsMarkov.exists_policy_eq_prodMkLeft.choose_spec n + +lemma policy_apply_eq_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] (n : β„•) + (h : Hist π“ž 𝓐 𝓨 n) (o : π“ž) : + alg.policy n (h, o) = alg.policyCondObs o := by + rw [policy_eq_prodMkLeft_policyCondObs, Kernel.prodMkLeft_apply] + +/-- The policy of a Markov algorithm at time `0` is a Markov kernel, hence so is +`alg.policyCondObs`. -/ +instance (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] : IsMarkovKernel alg.policyCondObs where + isProbabilityMeasure o := by + rw [← policy_apply_eq_policyCondObs alg 0 default o] + infer_instance + +lemma p0_eq_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] : + alg.p0 = alg.policyCondObs := by + ext o : 1 + rw [p0_apply, policy_apply_eq_policyCondObs] + +/-- The kernel `alg.policyCondObs` is determined by the policy at time `0`, since the empty history +is unique. -/ +lemma policyCondObs_eq_of_policy_zero_eq (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] {ΞΊ : Kernel π“ž 𝓐} + (h : alg.policy 0 = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 0)) : + alg.policyCondObs = ΞΊ := by + rw [policy_eq_prodMkLeft_policyCondObs, Kernel.prodMkLeft_inj] at h + exact h + +namespace IsMarkov + +variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} + {alg : Algorithm π“ž 𝓐 𝓨} {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} [IsFiniteMeasure P] + {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} {n N : β„•} + +lemma hasCondDistrib_action_history_obs [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + HasCondDistrib (A n) (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) + (alg.policyCondObs.prodMkLeft (Hist π“ž 𝓐 𝓨 n)) P := by + rw [← alg.policy_eq_prodMkLeft_policyCondObs] + exact h.hasCondDistrib_action n + +lemma hasCondDistrib_action_history_obs_of_isAlgEnvSeqUntil [alg.IsMarkov] + (h : IsAlgEnvSeqUntil O A Y alg env P N) (hn : n < N) : + HasCondDistrib (A n) (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) + (alg.policyCondObs.prodMkLeft (Hist π“ž 𝓐 𝓨 n)) P := by + rw [← alg.policy_eq_prodMkLeft_policyCondObs] + exact h.hasCondDistrib_action n hn + +/-- The conditional distribution of the action at time `n` given the observation at time `n` is +`alg.policyCondObs`. -/ +lemma hasCondDistrib_action [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + HasCondDistrib (A n) (O n) alg.policyCondObs P := + (hasCondDistrib_action_history_obs h n).comp_right + +/-- The conditional distribution of the action at time `n < N` given the observation at time `n` +is `alg.policyCondObs`. -/ +lemma hasCondDistrib_action_of_isAlgEnvSeqUntil [alg.IsMarkov] + (h : IsAlgEnvSeqUntil O A Y alg env P N) (hn : n < N) : + HasCondDistrib (A n) (O n) alg.policyCondObs P := + (hasCondDistrib_action_history_obs_of_isAlgEnvSeqUntil h hn).comp_right + +/-- The law of the action at time `n` is the law of the observation at time `n` composed with +`alg.policyCondObs`. -/ +lemma hasLaw_action_comp [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + HasLaw (A n) (alg.policyCondObs βˆ˜β‚˜ (P.map (O n))) P := + (hasCondDistrib_action h n).hasLaw_comp + +/-- Conditionally on an event determined by the history before time `n` and the observation at +time `n`, on which that observation is equal to `b`, the action at time `n` has law +`alg.policyCondObs b`. -/ +lemma hasLaw_action_cond [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) + {s : Set (Hist π“ž 𝓐 𝓨 n Γ— π“ž)} (hs : MeasurableSet s) {b : π“ž} (hsb : βˆ€ u ∈ s, u.2 = b) + (hP : P ((fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s) β‰  0) : + HasLaw (A n) (alg.policyCondObs b) P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s] := by + refine (hasCondDistrib_action_history_obs h n).hasLaw_cond (h.measurable_action _) hs + (fun u hu ↦ ?_) hP + rw [Kernel.prodMkLeft_apply, hsb u hu] + +/-- Conditionally on an event determined by the history before time `n` and the observation at +time `n`, on which that observation is constant, the action at time `n` is independent of the +history before time `n` and of the observation at time `n`. -/ +lemma indepFun_history_obs_action_cond [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) + {s : Set (Hist π“ž 𝓐 𝓨 n Γ— π“ž)} (hs : MeasurableSet s) {b : π“ž} (hsb : βˆ€ u ∈ s, u.2 = b) : + (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) + βŸ‚α΅’[P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s]] A n := by + have hO := h.measurable_obs + have hA := h.measurable_action + have hY := h.measurable_feedback + refine (hasCondDistrib_action_history_obs h n).indepFun_cond (by fun_prop) hs + (Ξ· := alg.policyCondObs b) fun u hu ↦ ?_ + rw [Kernel.prodMkLeft_apply, hsb u hu] + +variable [StandardBorelSpace π“ž] [Nonempty π“ž] [StandardBorelSpace 𝓐] [Nonempty 𝓐] + [StandardBorelSpace 𝓨] [Nonempty 𝓨] + +/-- The action at time `n` is conditionally independent of the history before time `n`, given the +observation at time `n`. -/ +lemma condIndepFun_action_history [StandardBorelSpace Ξ©] [alg.IsMarkov] + (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + A n βŸ‚α΅’[O n, h.measurable_obs n; P] history O A Y n := by + have hO := h.measurable_obs + have hA := h.measurable_action + have hY := h.measurable_feedback + refine condIndepFun_of_exists_condDistrib_prod_ae_eq_prodMkLeft (Ξ· := alg.policyCondObs) + (by fun_prop) (by fun_prop) (by fun_prop) ?_ + exact HasCondDistrib.condDistrib_eq (hasCondDistrib_action_history_obs h n) + +/-- The action at time `n` is conditionally independent of the history before time `n` and the +observation at time `n`, given the observation at time `n`. -/ +lemma condIndepFun_action_history_obs [StandardBorelSpace Ξ©] [alg.IsMarkov] + (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + A n βŸ‚α΅’[O n, h.measurable_obs n; P] (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) := by + have hO := h.measurable_obs + have hA := h.measurable_action + have hY := h.measurable_feedback + exact (condIndepFun_action_history h n).prod_right (by fun_prop) (by fun_prop) (by fun_prop) + +end IsMarkov + +section Markov + +variable {ΞΊ : Kernel π“ž 𝓐} [IsMarkovKernel ΞΊ] + +/-- An algorithm with a policy that depends only on the current observation. -/ +def markov (ΞΊ : Kernel π“ž 𝓐) [IsMarkovKernel ΞΊ] : Algorithm π“ž 𝓐 𝓨 where + policy n := ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n) + +@[simp] +lemma policy_markov (n : β„•) : + (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policy n = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n) := rfl + +@[simp] +lemma p0_markov : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).p0 = ΞΊ := by + ext o : 1 + rw [p0_apply, policy_markov, Kernel.prodMkLeft_apply] + +instance : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).IsMarkov where + exists_policy_eq_prodMkLeft := ⟨κ, fun _ ↦ rfl⟩ + +@[simp] +lemma policyCondObs_markov : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policyCondObs = ΞΊ := + policyCondObs_eq_of_policy_zero_eq _ rfl + +end Markov + +end Algorithm + +namespace IsAlgEnvSeq + +variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} + {ΞΊ : Kernel π“ž 𝓐} [IsMarkovKernel ΞΊ] {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} [IsFiniteMeasure P] + {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} + +/-- The conditional distribution of the action at time `n` given the observation at time `n` +is `ΞΊ`. -/ +lemma hasCondDistrib_action_markov (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) : + HasCondDistrib (A n) (O n) ΞΊ P := by + simpa using Algorithm.IsMarkov.hasCondDistrib_action h n + +/-- The conditional distribution of the action at time `n` given the observation at time `n` +is `ΞΊ`. -/ +lemma condDistrib_action_markov [StandardBorelSpace 𝓐] [Nonempty 𝓐] + (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) : + condDistrib (A n) (O n) P =ᡐ[P.map (O n)] ΞΊ := + (hasCondDistrib_action_markov h n).condDistrib_eq + +/-- Conditionally on an event determined by the history before time `n` and the observation at +time `n`, on which that observation is equal to `b`, the action at time `n` has law `ΞΊ b`. -/ +lemma hasLaw_action_cond_markov (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) + {s : Set (Hist π“ž 𝓐 𝓨 n Γ— π“ž)} (hs : MeasurableSet s) {b : π“ž} (hsb : βˆ€ u ∈ s, u.2 = b) + (hP : P ((fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s) β‰  0) : + HasLaw (A n) (ΞΊ b) P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s] := by + simpa using Algorithm.IsMarkov.hasLaw_action_cond h n hs hsb hP + +/-- Conditionally on an event determined by the history before time `n` and the observation at +time `n`, on which that observation is constant, the action at time `n` is independent of the +history before time `n` and of the observation at time `n`. -/ +lemma indepFun_history_obs_action_cond_markov + (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) + {s : Set (Hist π“ž 𝓐 𝓨 n Γ— π“ž)} (hs : MeasurableSet s) {b : π“ž} (hsb : βˆ€ u ∈ s, u.2 = b) : + (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) + βŸ‚α΅’[P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s]] A n := + Algorithm.IsMarkov.indepFun_history_obs_action_cond h n hs hsb + +/-- The action at time `n` is conditionally independent of the history before time `n`, given the +observation at time `n`. -/ +lemma condIndepFun_action_history_markov [StandardBorelSpace Ξ©] + [StandardBorelSpace π“ž] [Nonempty π“ž] [StandardBorelSpace 𝓐] [Nonempty 𝓐] + [StandardBorelSpace 𝓨] [Nonempty 𝓨] + (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) : + A n βŸ‚α΅’[O n, h.measurable_obs n; P] history O A Y n := + Algorithm.IsMarkov.condIndepFun_action_history h n + +end IsAlgEnvSeq + +end Learning diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean index 7ade0abe..653dafcb 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean @@ -5,7 +5,7 @@ Authors: GaΓ«tan SerrΓ© -/ module -public import LeanMachineLearning.SequentialLearning.Algorithm +public import LeanMachineLearning.SequentialLearning.Algorithms.Markov import LeanMachineLearning.ForMathlib.Probability.Independence.IndepFun @@ -47,16 +47,28 @@ noncomputable def randomSampling (ΞΌ : Measure 𝓐) [IsProbabilityMeasure ΞΌ] : Algorithm π“ž 𝓐 𝓨 where policy _ := Kernel.const _ ΞΌ +/-- The random sampling algorithm is the Markov algorithm with the constant kernel `ΞΌ`. -/ +lemma randomSampling_eq_markov : + (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨) = Algorithm.markov (Kernel.const π“ž ΞΌ) := rfl + +instance : (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).IsMarkov where + exists_policy_eq_prodMkLeft := ⟨Kernel.const π“ž ΞΌ, fun _ ↦ rfl⟩ + +@[simp] +lemma policyCondObs_randomSampling : + (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).policyCondObs = Kernel.const π“ž ΞΌ := + Algorithm.policyCondObs_eq_of_policy_zero_eq _ rfl + namespace randomSampling variable {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} {env : Environment π“ž 𝓐 𝓨} -/-- Each action follows the distribution ΞΌ. -/ +/-- Each action of the random sampling algorithm follows the distribution ΞΌ. -/ lemma hasLaw_action (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) env P) (n : β„•) : HasLaw (A n) ΞΌ P := (h.hasCondDistrib_action n).hasLaw_of_const -/-- Actions are mutually independent. -/ +/-- Actions of the random sampling algorithm are mutually independent. -/ lemma iIndep_action (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) env P) : iIndepFun A P := by have hO := h.measurable_obs diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean b/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean index 7c6f3c4d..0ae9e703 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean @@ -40,6 +40,9 @@ noncomputable def uniformAlgorithm [Finite 𝓐] [Nonempty 𝓐] : Algorithm π“ž 𝓐 𝓨 := randomSampling (uniformOn Set.univ) +instance [Finite 𝓐] [Nonempty 𝓐] : (uniformAlgorithm : Algorithm π“ž 𝓐 𝓨).IsMarkov := + inferInstanceAs (randomSampling (uniformOn Set.univ) : Algorithm π“ž 𝓐 𝓨).IsMarkov + lemma absolutelyContinuous_uniformAlgorithm [Finite 𝓐] [Nonempty 𝓐] {alg : Algorithm π“ž 𝓐 𝓨} : alg β‰ͺₐ uniformAlgorithm where policy n h := Measure.absolutelyContinuous_of_measure_singleton_ne_zero From fe465e321dac6e228be0b44cf446107372c60bd2 Mon Sep 17 00:00:00 2001 From: Remy Degenne Date: Thu, 17 Sep 2026 16:48:51 +0200 Subject: [PATCH 4/7] make it time-dependent --- .../SequentialLearning/Algorithms/Markov.lean | 140 ++++++++++-------- .../Algorithms/RandomSampling/Basic.lean | 15 +- 2 files changed, 84 insertions(+), 71 deletions(-) diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean b/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean index f8872080..00fd20e9 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean @@ -9,27 +9,28 @@ public import LeanMachineLearning.ForMathlib.Probability.Kernel.Composition.MapC public import LeanMachineLearning.SequentialLearning.Algorithm /-! -# Algorithms that depend only on the observation +# Algorithms that depend only on the time and the observation In general, an algorithm `alg : Algorithm π“ž 𝓐 𝓨` has a policy `alg.policy n : Kernel (Hist π“ž 𝓐 𝓨 n Γ— π“ž) 𝓐`, which depends on the history before time `n` and on -the observation at time `n`. In some cases, the policy depends only on the observation. -We say that `alg` is a *Markov algorithm* if there exists a Markov kernel `ΞΊ : Kernel π“ž 𝓐` such that -`alg.policy n = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n)` for all `n`. +the observation at time `n`. In some cases, the policy depends only on the time and on the +observation. We say that `alg` is a *Markov algorithm* if there exists a sequence of Markov kernels +`ΞΊ : β„• β†’ Kernel π“ž 𝓐` such that `alg.policy n = (ΞΊ n).prodMkLeft (Hist π“ž 𝓐 𝓨 n)` for all `n`. ## Main definitions -* `Algorithm.IsMarkov alg`: the policy of `alg` depends only on the current observation, - not on the history and not on the time. -* `Algorithm.policyCondObs alg`: the kernel representing the conditional distribution of the - action given the observation in a Markov algorithm `alg`. -* `Algorithm.markov ΞΊ`: an algorithm with a policy that depends only on the current observation, - given by the Markov kernel `ΞΊ`. +* `Algorithm.IsMarkov alg`: the policy of `alg` at time `n` depends only on `n` and on the current + observation, not on the history. +* `Algorithm.policyCondObs alg n`: the kernel representing the conditional distribution of the + action at time `n` given the observation at time `n` in a Markov algorithm `alg`. +* `Algorithm.markov ΞΊ`: the algorithm whose action at time `n` is drawn from the Markov kernel + `ΞΊ n` applied to the observation at time `n`. ## Main statements * `Algorithm.IsMarkov.hasCondDistrib_action`: in a run of a Markov algorithm, the conditional - distribution of the action at time `n` given the observation at time `n` is `alg.policyCondObs`. + distribution of the action at time `n` given the observation at time `n` is + `alg.policyCondObs n`. * `Algorithm.IsMarkov.condIndepFun_action_history`: in a run of a Markov algorithm, the action at time `n` is conditionally independent of the history before time `n` given the observation at time `n`. @@ -47,45 +48,46 @@ namespace Learning variable {π“ž 𝓐 𝓨 : Type*} {mπ“ž : MeasurableSpace π“ž} {m𝓐 : MeasurableSpace 𝓐} {m𝓨 : MeasurableSpace 𝓨} -/-- The policy of the algorithm depends only on the current observation, not on the history and -not on the time. -/ +/-- The policy of the algorithm at time `n` depends only on `n` and on the current observation, +not on the history. -/ class Algorithm.IsMarkov (alg : Algorithm π“ž 𝓐 𝓨) : Prop where - exists_policy_eq_prodMkLeft : βˆƒ ΞΊ : Kernel π“ž 𝓐, βˆ€ n, alg.policy n = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n) + exists_policy_eq_prodMkLeft : βˆƒ ΞΊ : β„• β†’ Kernel π“ž 𝓐, (βˆ€ n, IsMarkovKernel (ΞΊ n)) ∧ + βˆ€ n, alg.policy n = (ΞΊ n).prodMkLeft (Hist π“ž 𝓐 𝓨 n) namespace Algorithm -/-- The kernel representing the conditional distribution of the action given the observation -in a Markov algorithm. -/ +/-- The kernel representing the conditional distribution of the action at time `n` given the +observation at time `n` in a Markov algorithm. -/ noncomputable -def policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [h_markov : alg.IsMarkov] : Kernel π“ž 𝓐 := - h_markov.exists_policy_eq_prodMkLeft.choose +def policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [h_markov : alg.IsMarkov] (n : β„•) : Kernel π“ž 𝓐 := + h_markov.exists_policy_eq_prodMkLeft.choose n + +instance (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] (n : β„•) : IsMarkovKernel (alg.policyCondObs n) := + IsMarkov.exists_policy_eq_prodMkLeft.choose_spec.1 n lemma policy_eq_prodMkLeft_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] (n : β„•) : - alg.policy n = alg.policyCondObs.prodMkLeft (Hist π“ž 𝓐 𝓨 n) := - IsMarkov.exists_policy_eq_prodMkLeft.choose_spec n + alg.policy n = (alg.policyCondObs n).prodMkLeft (Hist π“ž 𝓐 𝓨 n) := + IsMarkov.exists_policy_eq_prodMkLeft.choose_spec.2 n lemma policy_apply_eq_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] (n : β„•) (h : Hist π“ž 𝓐 𝓨 n) (o : π“ž) : - alg.policy n (h, o) = alg.policyCondObs o := by + alg.policy n (h, o) = alg.policyCondObs n o := by rw [policy_eq_prodMkLeft_policyCondObs, Kernel.prodMkLeft_apply] -/-- The policy of a Markov algorithm at time `0` is a Markov kernel, hence so is -`alg.policyCondObs`. -/ -instance (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] : IsMarkovKernel alg.policyCondObs where - isProbabilityMeasure o := by - rw [← policy_apply_eq_policyCondObs alg 0 default o] - infer_instance - lemma p0_eq_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] : - alg.p0 = alg.policyCondObs := by + alg.p0 = alg.policyCondObs 0 := by ext o : 1 rw [p0_apply, policy_apply_eq_policyCondObs] -/-- The kernel `alg.policyCondObs` is determined by the policy at time `0`, since the empty history -is unique. -/ -lemma policyCondObs_eq_of_policy_zero_eq (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] {ΞΊ : Kernel π“ž 𝓐} - (h : alg.policy 0 = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 0)) : - alg.policyCondObs = ΞΊ := by +/-- The kernel `alg.policyCondObs n` is determined by the policy at time `n`. The assumption +`Nonempty 𝓨` ensures that there are histories of every length. -/ +lemma policyCondObs_eq_of_policy_eq [Nonempty 𝓨] (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] {n : β„•} + {ΞΊ : Kernel π“ž 𝓐} (h : alg.policy n = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n)) : + alg.policyCondObs n = ΞΊ := by + rcases isEmpty_or_nonempty π“ž with hπ“ž | hπ“ž + Β· ext o : 1 + exact hπ“ž.elim o + have : Nonempty 𝓐 := Measure.nonempty_of_neZero (alg.policyCondObs n (Classical.arbitrary π“ž)) rw [policy_eq_prodMkLeft_policyCondObs, Kernel.prodMkLeft_inj] at h exact h @@ -97,43 +99,43 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} lemma hasCondDistrib_action_history_obs [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : HasCondDistrib (A n) (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) - (alg.policyCondObs.prodMkLeft (Hist π“ž 𝓐 𝓨 n)) P := by + ((alg.policyCondObs n).prodMkLeft (Hist π“ž 𝓐 𝓨 n)) P := by rw [← alg.policy_eq_prodMkLeft_policyCondObs] exact h.hasCondDistrib_action n lemma hasCondDistrib_action_history_obs_of_isAlgEnvSeqUntil [alg.IsMarkov] (h : IsAlgEnvSeqUntil O A Y alg env P N) (hn : n < N) : HasCondDistrib (A n) (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) - (alg.policyCondObs.prodMkLeft (Hist π“ž 𝓐 𝓨 n)) P := by + ((alg.policyCondObs n).prodMkLeft (Hist π“ž 𝓐 𝓨 n)) P := by rw [← alg.policy_eq_prodMkLeft_policyCondObs] exact h.hasCondDistrib_action n hn /-- The conditional distribution of the action at time `n` given the observation at time `n` is -`alg.policyCondObs`. -/ +`alg.policyCondObs n`. -/ lemma hasCondDistrib_action [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : - HasCondDistrib (A n) (O n) alg.policyCondObs P := + HasCondDistrib (A n) (O n) (alg.policyCondObs n) P := (hasCondDistrib_action_history_obs h n).comp_right /-- The conditional distribution of the action at time `n < N` given the observation at time `n` -is `alg.policyCondObs`. -/ +is `alg.policyCondObs n`. -/ lemma hasCondDistrib_action_of_isAlgEnvSeqUntil [alg.IsMarkov] (h : IsAlgEnvSeqUntil O A Y alg env P N) (hn : n < N) : - HasCondDistrib (A n) (O n) alg.policyCondObs P := + HasCondDistrib (A n) (O n) (alg.policyCondObs n) P := (hasCondDistrib_action_history_obs_of_isAlgEnvSeqUntil h hn).comp_right /-- The law of the action at time `n` is the law of the observation at time `n` composed with -`alg.policyCondObs`. -/ +`alg.policyCondObs n`. -/ lemma hasLaw_action_comp [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : - HasLaw (A n) (alg.policyCondObs βˆ˜β‚˜ (P.map (O n))) P := + HasLaw (A n) (alg.policyCondObs n βˆ˜β‚˜ (P.map (O n))) P := (hasCondDistrib_action h n).hasLaw_comp /-- Conditionally on an event determined by the history before time `n` and the observation at time `n`, on which that observation is equal to `b`, the action at time `n` has law -`alg.policyCondObs b`. -/ +`alg.policyCondObs n b`. -/ lemma hasLaw_action_cond [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) {s : Set (Hist π“ž 𝓐 𝓨 n Γ— π“ž)} (hs : MeasurableSet s) {b : π“ž} (hsb : βˆ€ u ∈ s, u.2 = b) (hP : P ((fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s) β‰  0) : - HasLaw (A n) (alg.policyCondObs b) P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s] := by + HasLaw (A n) (alg.policyCondObs n b) P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s] := by refine (hasCondDistrib_action_history_obs h n).hasLaw_cond (h.measurable_action _) hs (fun u hu ↦ ?_) hP rw [Kernel.prodMkLeft_apply, hsb u hu] @@ -149,7 +151,7 @@ lemma indepFun_history_obs_action_cond [alg.IsMarkov] (h : IsAlgEnvSeq O A Y alg have hA := h.measurable_action have hY := h.measurable_feedback refine (hasCondDistrib_action_history_obs h n).indepFun_cond (by fun_prop) hs - (Ξ· := alg.policyCondObs b) fun u hu ↦ ?_ + (Ξ· := alg.policyCondObs n b) fun u hu ↦ ?_ rw [Kernel.prodMkLeft_apply, hsb u hu] variable [StandardBorelSpace π“ž] [Nonempty π“ž] [StandardBorelSpace 𝓐] [Nonempty 𝓐] @@ -163,7 +165,7 @@ lemma condIndepFun_action_history [StandardBorelSpace Ξ©] [alg.IsMarkov] have hO := h.measurable_obs have hA := h.measurable_action have hY := h.measurable_feedback - refine condIndepFun_of_exists_condDistrib_prod_ae_eq_prodMkLeft (Ξ· := alg.policyCondObs) + refine condIndepFun_of_exists_condDistrib_prod_ae_eq_prodMkLeft (Ξ· := alg.policyCondObs n) (by fun_prop) (by fun_prop) (by fun_prop) ?_ exact HasCondDistrib.condDistrib_eq (hasCondDistrib_action_history_obs h n) @@ -181,27 +183,29 @@ end IsMarkov section Markov -variable {ΞΊ : Kernel π“ž 𝓐} [IsMarkovKernel ΞΊ] +variable {ΞΊ : β„• β†’ Kernel π“ž 𝓐} [βˆ€ n, IsMarkovKernel (ΞΊ n)] -/-- An algorithm with a policy that depends only on the current observation. -/ -def markov (ΞΊ : Kernel π“ž 𝓐) [IsMarkovKernel ΞΊ] : Algorithm π“ž 𝓐 𝓨 where - policy n := ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n) +/-- The algorithm whose action at time `n` is drawn from `ΞΊ n` applied to the observation at +time `n`, whatever the history. -/ +def markov (ΞΊ : β„• β†’ Kernel π“ž 𝓐) [βˆ€ n, IsMarkovKernel (ΞΊ n)] : Algorithm π“ž 𝓐 𝓨 where + policy n := (ΞΊ n).prodMkLeft (Hist π“ž 𝓐 𝓨 n) @[simp] lemma policy_markov (n : β„•) : - (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policy n = ΞΊ.prodMkLeft (Hist π“ž 𝓐 𝓨 n) := rfl + (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policy n = (ΞΊ n).prodMkLeft (Hist π“ž 𝓐 𝓨 n) := rfl @[simp] -lemma p0_markov : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).p0 = ΞΊ := by +lemma p0_markov : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).p0 = ΞΊ 0 := by ext o : 1 rw [p0_apply, policy_markov, Kernel.prodMkLeft_apply] instance : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).IsMarkov where - exists_policy_eq_prodMkLeft := ⟨κ, fun _ ↦ rfl⟩ + exists_policy_eq_prodMkLeft := ⟨κ, inferInstance, fun _ ↦ rfl⟩ @[simp] -lemma policyCondObs_markov : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policyCondObs = ΞΊ := - policyCondObs_eq_of_policy_zero_eq _ rfl +lemma policyCondObs_markov [Nonempty 𝓨] (n : β„•) : + (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policyCondObs n = ΞΊ n := + policyCondObs_eq_of_policy_eq _ rfl end Markov @@ -210,29 +214,37 @@ end Algorithm namespace IsAlgEnvSeq variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} - {ΞΊ : Kernel π“ž 𝓐} [IsMarkovKernel ΞΊ] {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} [IsFiniteMeasure P] + {ΞΊ : β„• β†’ Kernel π“ž 𝓐} [βˆ€ n, IsMarkovKernel (ΞΊ n)] {env : Environment π“ž 𝓐 𝓨} + {P : Measure Ξ©} [IsFiniteMeasure P] {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} +lemma hasCondDistrib_action_history_obs_markov + (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) : + HasCondDistrib (A n) (fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ((ΞΊ n).prodMkLeft (Hist π“ž 𝓐 𝓨 n)) P := + h.hasCondDistrib_action n + /-- The conditional distribution of the action at time `n` given the observation at time `n` -is `ΞΊ`. -/ +is `ΞΊ n`. -/ lemma hasCondDistrib_action_markov (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) : - HasCondDistrib (A n) (O n) ΞΊ P := by - simpa using Algorithm.IsMarkov.hasCondDistrib_action h n + HasCondDistrib (A n) (O n) (ΞΊ n) P := + (hasCondDistrib_action_history_obs_markov h n).comp_right /-- The conditional distribution of the action at time `n` given the observation at time `n` -is `ΞΊ`. -/ +is `ΞΊ n`. -/ lemma condDistrib_action_markov [StandardBorelSpace 𝓐] [Nonempty 𝓐] (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) : - condDistrib (A n) (O n) P =ᡐ[P.map (O n)] ΞΊ := + condDistrib (A n) (O n) P =ᡐ[P.map (O n)] ΞΊ n := (hasCondDistrib_action_markov h n).condDistrib_eq /-- Conditionally on an event determined by the history before time `n` and the observation at -time `n`, on which that observation is equal to `b`, the action at time `n` has law `ΞΊ b`. -/ +time `n`, on which that observation is equal to `b`, the action at time `n` has law `ΞΊ n b`. -/ lemma hasLaw_action_cond_markov (h : IsAlgEnvSeq O A Y (Algorithm.markov ΞΊ) env P) (n : β„•) {s : Set (Hist π“ž 𝓐 𝓨 n Γ— π“ž)} (hs : MeasurableSet s) {b : π“ž} (hsb : βˆ€ u ∈ s, u.2 = b) (hP : P ((fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s) β‰  0) : - HasLaw (A n) (ΞΊ b) P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s] := by - simpa using Algorithm.IsMarkov.hasLaw_action_cond h n hs hsb hP + HasLaw (A n) (ΞΊ n b) P[|(fun Ο‰ ↦ (history O A Y n Ο‰, O n Ο‰)) ⁻¹' s] := by + refine (hasCondDistrib_action_history_obs_markov h n).hasLaw_cond (h.measurable_action _) hs + (fun u hu ↦ ?_) hP + rw [Kernel.prodMkLeft_apply, hsb u hu] /-- Conditionally on an event determined by the history before time `n` and the observation at time `n`, on which that observation is constant, the action at time `n` is independent of the diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean index 653dafcb..4348030e 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean @@ -47,17 +47,18 @@ noncomputable def randomSampling (ΞΌ : Measure 𝓐) [IsProbabilityMeasure ΞΌ] : Algorithm π“ž 𝓐 𝓨 where policy _ := Kernel.const _ ΞΌ -/-- The random sampling algorithm is the Markov algorithm with the constant kernel `ΞΌ`. -/ +/-- The random sampling algorithm is the Markov algorithm with the constant kernel `ΞΌ` at every +time. -/ lemma randomSampling_eq_markov : - (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨) = Algorithm.markov (Kernel.const π“ž ΞΌ) := rfl + (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨) = Algorithm.markov (fun _ ↦ Kernel.const π“ž ΞΌ) := rfl -instance : (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).IsMarkov where - exists_policy_eq_prodMkLeft := ⟨Kernel.const π“ž ΞΌ, fun _ ↦ rfl⟩ +instance : (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).IsMarkov := + inferInstanceAs (Algorithm.markov (fun _ ↦ Kernel.const π“ž ΞΌ) : Algorithm π“ž 𝓐 𝓨).IsMarkov @[simp] -lemma policyCondObs_randomSampling : - (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).policyCondObs = Kernel.const π“ž ΞΌ := - Algorithm.policyCondObs_eq_of_policy_zero_eq _ rfl +lemma policyCondObs_randomSampling [Nonempty 𝓨] (n : β„•) : + (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).policyCondObs n = Kernel.const π“ž ΞΌ := + Algorithm.policyCondObs_eq_of_policy_eq _ rfl namespace randomSampling From 276a5d30363b25536b17910cc300e68fd13acae2 Mon Sep 17 00:00:00 2001 From: Remy Degenne Date: Sat, 3 Oct 2026 14:56:27 +0200 Subject: [PATCH 5/7] rename more --- LMLTutorial/Pages/DefiningAlgorithm.lean | 22 +- .../Online/Bandit/Algorithms/ETC.lean | 6 +- .../Online/Bandit/Algorithms/TS.lean | 6 +- .../Online/Bandit/Algorithms/UCB.lean | 6 +- .../Online/Bandit/ArrayProbSpace.lean | 9 +- .../SequentialLearning/Algorithm.lean | 51 ++-- .../SequentialLearning/AlgorithmDensity.lean | 4 +- .../SequentialLearning/Algorithms/Markov.lean | 10 +- .../Algorithms/RandomSampling/Basic.lean | 24 +- .../Algorithms/RandomSampling/Tendsto.lean | 45 +-- .../Algorithms/RoundRobin.lean | 4 +- .../Algorithms/Uniform.lean | 4 +- .../BayesStationaryEnv.lean | 10 +- .../SequentialLearning/Comap.lean | 67 ++--- .../SequentialLearning/Deterministic.lean | 268 +++++++++--------- .../SequentialLearning/EvaluationEnv.lean | 107 +++---- .../IonescuTulceaSpace.lean | 9 +- .../SequentialLearning/Means.lean | 12 +- .../SequentialLearning/README.md | 22 +- .../SequentialLearning/StationaryEnv.lean | 206 +++++++------- 20 files changed, 460 insertions(+), 432 deletions(-) diff --git a/LMLTutorial/Pages/DefiningAlgorithm.lean b/LMLTutorial/Pages/DefiningAlgorithm.lean index 7f781a5c..5fb570ed 100644 --- a/LMLTutorial/Pages/DefiningAlgorithm.lean +++ b/LMLTutorial/Pages/DefiningAlgorithm.lean @@ -61,14 +61,14 @@ The `policy` field contains for each time `n` a kernel from that history togethe That is, it maps every possible history and current observation to a random action at time `n` (and that map is measurable). The `isMarkovKernel_policy` field records that the measure describing the action is a probability measure (and it is in square brackets to tell Lean to infer it automatically whenever possible). At time `0` the history is empty: `Hist π“ž 𝓐 𝓨 0` has a unique element, and the distribution of the first action given the first observation is `policy 0` applied to that element. -That kernel is called `Algorithm.p0`. +That kernel is called `Algorithm.policyZero`. Many settings have no observations at all: the algorithm sees only the past rounds. Those are described by taking `π“ž = Unit`, and we write `noObs Ξ©` for the corresponding (constant) observation process. -If the algorithms actions are not random, we can use the `detAlgorithm` definition to build an algorithm from the data of a measurable function for the action at each time, as a function of the history before that time and of the current observation. +If the algorithms actions are not random, we can use the `Algorithm.deterministic` definition to build an algorithm from the data of a measurable function for the action at each time, as a function of the history before that time and of the current observation. The first action is the value of that function at time `0` on the empty history. -{docstring detAlgorithm} +{docstring Algorithm.deterministic} We can see here that we did not need to prove that the kernels are `IsMarkovKernel`. Lean knows that deterministic kernels are Markov. @@ -79,17 +79,17 @@ The `Environment` structure is the mirror of the `Algorithm` structure, with a k `obs n` gives the distribution of the observation at time `n` given the history before `n`. `feedback n` gives the distribution of the feedback at time `n` given the history before `n`, the observation and the action at time `n`. -The distribution of the first observation is `obs 0` applied to the empty history; it is called `Environment.obs0`. -The distribution of the first feedback given the first observation and action is `feedback 0` applied to the empty history; it is called `Environment.Ξ½0`. +The distribution of the first observation is `obs 0` applied to the empty history; it is called `Environment.obsZero`. +The distribution of the first feedback given the first observation and action is `feedback 0` applied to the empty history; it is called `Environment.feedbackZero`. In many applications neither the observation nor the feedback depends on the prior history: the observation at time `n` has a fixed law, and the feedback depends only on the current observation and action. -We provide an `obliviousEnv` definition that builds an environment for those cases from a sequence of observation laws and a sequence of feedback kernels. +We provide an `Environment.oblivious` definition that builds an environment for those cases from a sequence of observation laws and a sequence of feedback kernels. -{docstring obliviousEnv} +{docstring Environment.oblivious} -If furthermore those sequences do not change with time, we can use the `stationaryEnv` definition to build the environment. +If furthermore those sequences do not change with time, we can use the `Environment.stationary` definition to build the environment. -{docstring stationaryEnv} +{docstring Environment.stationary} When there is no observation (`π“ž = Unit`), the feedback depends only on the last action. `Environment.banditSeq` builds such an environment from a sequence of kernels `Ξ½ : β„• β†’ Kernel 𝓐 𝓨`, and `Environment.bandit` from a single kernel `Ξ½ : Kernel 𝓐 𝓨` used at every time. @@ -129,7 +129,7 @@ The environment is thus simply `Environment.bandit Ξ½` for some kernel `Ξ½ : Ker The UCB algorithm chooses at time `n` the action that maximizes the sum of the empirical mean reward and an exploration bonus. It starts by choosing each action once and then chooses $`\arg\max_a (\hat{\mu}_{n,a} + \sqrt{\frac{2c \log (n + 1)}{N_{n,a}}})`, in which $`\hat{\mu}_{n,a}` is the empirical mean reward of action `a` before time `n` (`empMean'` in the code), $`N_{n,a}` is the number of times action `a` has been chosen before time `n` (`pullCount'` in the code), and `c` is a parameter of the algorithm. -To define the algorithm, we first define the exploration bonus and the next action function, and then we use `detAlgorithm` to build the algorithm. +To define the algorithm, we first define the exploration bonus and the next action function, and then we use `Algorithm.deterministic` to build the algorithm. We also need to prove that the next action function is measurable, which is done by the `measurable_nextArm` lemma. Note that we are careful to use a measurable version of the argmax function, `argmax`. @@ -141,7 +141,7 @@ Note that we are careful to use a measurable version of the argmax function, `ar {docstring Bandits.ucbAlgorithm} -The last line builds the algorithm using `detAlgorithm` and the function `UCB.nextArm`. +The last line builds the algorithm using `Algorithm.deterministic` and the function `UCB.nextArm`. Its measurability is proved by the `fun_prop` tactic, which proves measurability of functions by using lemmas tagged with `@[fun_prop]`. The first action of the algorithm is `UCB.nextArm K c 0` applied to the empty history, which is 0 as an element of `Fin K`. diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean b/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean index 3e014b73..6a28cc8f 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/ETC.lean @@ -48,7 +48,7 @@ variable (K) in to `ETC.nextArm`. -/ noncomputable def etcAlgorithm [NeZero K] (m : β„•) : Algorithm Unit (Fin K) ℝ := - detAlgorithm (fun n p ↦ ETC.nextArm K m n p.1) (by fun_prop) + Algorithm.deterministic (fun n p ↦ ETC.nextArm K m n p.1) (by fun_prop) end AlgorithmDefinition @@ -64,7 +64,7 @@ lemma isAlgEnvSeqUntil_roundRobinAlgorithm (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) : IsAlgEnvSeqUntil O A R (roundRobinAlgorithm K) (Environment.bandit Ξ½) P (K * m) := by refine h.isAlgEnvSeqUntil_of_policy_eq fun n hn ↦ ?_ - simp only [roundRobinAlgorithm, detAlgorithm_policy, etcAlgorithm] + simp only [roundRobinAlgorithm, Algorithm.deterministic_policy, etcAlgorithm] congr 1 with p simp [ETC.nextArm, hn] @@ -73,7 +73,7 @@ section AlgorithmBehavior lemma arm_ae_eq_nextArm (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) (n : β„•) : A n =ᡐ[P] fun Ο‰ ↦ nextArm K m n (history O A R n Ο‰) := - h.action_detAlgorithm_ae_eq n + h.action_deterministic_ae_eq n /-- For `n < K * m`, the arm pulled at time `n` is the arm `n % K`. -/ lemma arm_of_lt (h : IsAlgEnvSeq O A R (etcAlgorithm K m) (Environment.bandit Ξ½) P) diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean b/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean index 4ffe5cb1..12a46b0f 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/TS.lean @@ -81,9 +81,9 @@ variable {P : Measure Ξ©} [IsProbabilityMeasure P] /-- The first action of Thompson sampling is sampled according to its probability of being optimal under the prior over environments. -/ -lemma TS.p0_tsAlgorithm : - (tsAlgorithm Q ΞΊ).p0 () = Q.map (bestAction ΞΊ id) := by - rw [Algorithm.p0_apply] +lemma TS.policyZero_tsAlgorithm : + (tsAlgorithm Q ΞΊ).policyZero () = Q.map (bestAction ΞΊ id) := by + rw [Algorithm.policyZero_apply] dsimp only [tsAlgorithm] rw [TS.policy, Kernel.prodMkRight_apply, Kernel.map_apply _ (by fun_prop), IT.bayesTrajMeasurePosterior_zero, Kernel.const_apply] diff --git a/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean b/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean index 2b84884f..38ca7461 100644 --- a/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean +++ b/LeanMachineLearning/Online/Bandit/Algorithms/UCB.lean @@ -54,7 +54,7 @@ variable (K) in /-- The UCB algorithm. -/ noncomputable def ucbAlgorithm [NeZero K] (c : ℝ) : Algorithm Unit (Fin K) ℝ := - detAlgorithm (fun n p ↦ UCB.nextArm K c n p.1) (by fun_prop) + Algorithm.deterministic (fun n p ↦ UCB.nextArm K c n p.1) (by fun_prop) end Algorithm namespace UCB @@ -70,7 +70,7 @@ lemma isAlgEnvSeqUntil_roundRobinAlgorithm (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) : IsAlgEnvSeqUntil O A R (roundRobinAlgorithm K) (Environment.bandit Ξ½) P K := by refine h.isAlgEnvSeqUntil_of_policy_eq fun n hn ↦ ?_ - simp only [roundRobinAlgorithm, detAlgorithm_policy, ucbAlgorithm] + simp only [roundRobinAlgorithm, Algorithm.deterministic_policy, ucbAlgorithm] congr 1 with p simp [UCB.nextArm, hn] @@ -100,7 +100,7 @@ lemma arm_zero (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) lemma arm_ae_eq_nextArm (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (n : β„•) : A n =ᡐ[P] fun Ο‰ ↦ nextArm K c n (history O A R n Ο‰) := - h.action_detAlgorithm_ae_eq n + h.action_deterministic_ae_eq n lemma ucbIndex_le_ucbIndex_arm (h : IsAlgEnvSeq O A R (ucbAlgorithm K c) (Environment.bandit Ξ½) P) (a : Fin K) (hn : K ≀ n) : diff --git a/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean b/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean index b5320351..55061bff 100644 --- a/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean +++ b/LeanMachineLearning/Online/Bandit/ArrayProbSpace.lean @@ -225,7 +225,7 @@ lemma algFunction_map (alg : Algorithm Unit 𝓐 𝓑) (n : β„•) (h : Hist Unit /-- The initial action is the image of a uniform random variable by `algFunction alg 0 default`. -/ lemma algFunction_zero_map (alg : Algorithm Unit 𝓐 𝓑) : - volume.map (algFunction alg 0 default) = alg.p0 () := + volume.map (algFunction alg 0 default) = alg.policyZero () := algFunction_map alg 0 default @[fun_prop] @@ -778,8 +778,9 @@ lemma isAlgEnvSeq_arrayMeasure (alg : Algorithm Unit 𝓐 𝓑) (Ξ½ : Kernel hasCondDistrib_feedback := hasCondDistrib_reward alg Ξ½ lemma hasLaw_action_zero (alg : Algorithm Unit 𝓐 𝓑) (Ξ½ : Kernel 𝓐 𝓑) [IsMarkovKernel Ξ½] : - HasLaw (action alg 0) (alg.p0 ()) (arrayMeasure Ξ½) := by - have h : HasCondDistrib (action alg 0) (fun _ : probSpace 𝓐 𝓑 ↦ ()) alg.p0 (arrayMeasure Ξ½) := + HasLaw (action alg 0) (alg.policyZero ()) (arrayMeasure Ξ½) := by + have h : HasCondDistrib (action alg 0) (fun _ : probSpace 𝓐 𝓑 ↦ ()) alg.policyZero + (arrayMeasure Ξ½) := (isAlgEnvSeq_arrayMeasure alg Ξ½).hasCondDistrib_action_zero exact h.hasLaw_of_const' @@ -787,7 +788,7 @@ lemma hasCondDistrib_reward_zero (alg : Algorithm Unit 𝓐 𝓑) (Ξ½ : Kernel [IsMarkovKernel Ξ½] : HasCondDistrib (reward alg 0) (action alg 0) Ξ½ (arrayMeasure Ξ½) := by have h := (isAlgEnvSeq_arrayMeasure alg Ξ½).hasCondDistrib_feedback_zero - rw [Ξ½0_bandit] at h + rw [feedbackZero_bandit] at h simpa using hasCondDistrib_prodMk_left_unique_iff.mp h end Laws diff --git a/LeanMachineLearning/SequentialLearning/Algorithm.lean b/LeanMachineLearning/SequentialLearning/Algorithm.lean index 4508c870..20159435 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithm.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithm.lean @@ -151,47 +151,48 @@ lemma measurable_environment_iff (f : Ξ© β†’ Environment π“ž 𝓐 𝓨) : /-- Distribution of the first observation: the observation kernel at time `0` applied to the empty history. -/ -def Environment.obs0 (env : Environment π“ž 𝓐 𝓨) : Measure π“ž := +def Environment.obsZero (env : Environment π“ž 𝓐 𝓨) : Measure π“ž := env.obs 0 default deriving IsProbabilityMeasure -lemma Environment.obs0_def (env : Environment π“ž 𝓐 𝓨) : env.obs0 = env.obs 0 default := rfl +lemma Environment.obsZero_def (env : Environment π“ž 𝓐 𝓨) : env.obsZero = env.obs 0 default := rfl lemma Environment.obs_zero (env : Environment π“ž 𝓐 𝓨) (h : Hist π“ž 𝓐 𝓨 0) : - env.obs 0 h = env.obs0 := by + env.obs 0 h = env.obsZero := by rw [Unique.eq_default h] rfl /-- Distribution of the first action given the first observation: the policy at time `0` applied to the empty history. -/ -noncomputable def Algorithm.p0 (alg : Algorithm π“ž 𝓐 𝓨) : Kernel π“ž 𝓐 := +noncomputable def Algorithm.policyZero (alg : Algorithm π“ž 𝓐 𝓨) : Kernel π“ž 𝓐 := (alg.policy 0).sectR default deriving IsMarkovKernel -lemma Algorithm.p0_def (alg : Algorithm π“ž 𝓐 𝓨) : alg.p0 = (alg.policy 0).sectR default := rfl +lemma Algorithm.policyZero_def (alg : Algorithm π“ž 𝓐 𝓨) : + alg.policyZero = (alg.policy 0).sectR default := rfl -lemma Algorithm.p0_apply (alg : Algorithm π“ž 𝓐 𝓨) (o : π“ž) : - alg.p0 o = alg.policy 0 (default, o) := rfl +lemma Algorithm.policyZero_apply (alg : Algorithm π“ž 𝓐 𝓨) (o : π“ž) : + alg.policyZero o = alg.policy 0 (default, o) := rfl lemma Algorithm.policy_zero (alg : Algorithm π“ž 𝓐 𝓨) (h : Hist π“ž 𝓐 𝓨 0) (o : π“ž) : - alg.policy 0 (h, o) = alg.p0 o := by + alg.policy 0 (h, o) = alg.policyZero o := by rw [Unique.eq_default h] rfl /-- Distribution of the first feedback given the first observation and action: the feedback kernel at time `0` applied to the empty history. -/ -noncomputable def Environment.Ξ½0 (env : Environment π“ž 𝓐 𝓨) : Kernel (π“ž Γ— 𝓐) 𝓨 := +noncomputable def Environment.feedbackZero (env : Environment π“ž 𝓐 𝓨) : Kernel (π“ž Γ— 𝓐) 𝓨 := (env.feedback 0).comap (fun p ↦ ((default, p.1), p.2)) (by fun_prop) deriving IsMarkovKernel -lemma Environment.Ξ½0_def (env : Environment π“ž 𝓐 𝓨) : - env.Ξ½0 = (env.feedback 0).comap (fun p ↦ ((default, p.1), p.2)) (by fun_prop) := rfl +lemma Environment.feedbackZero_def (env : Environment π“ž 𝓐 𝓨) : + env.feedbackZero = (env.feedback 0).comap (fun p ↦ ((default, p.1), p.2)) (by fun_prop) := rfl -lemma Environment.Ξ½0_apply (env : Environment π“ž 𝓐 𝓨) (o : π“ž) (a : 𝓐) : - env.Ξ½0 (o, a) = env.feedback 0 ((default, o), a) := rfl +lemma Environment.feedbackZero_apply (env : Environment π“ž 𝓐 𝓨) (o : π“ž) (a : 𝓐) : + env.feedbackZero (o, a) = env.feedback 0 ((default, o), a) := rfl lemma Environment.feedback_zero (env : Environment π“ž 𝓐 𝓨) (h : Hist π“ž 𝓐 𝓨 0) (o : π“ž) (a : 𝓐) : - env.feedback 0 ((h, o), a) = env.Ξ½0 (o, a) := by + env.feedback 0 ((h, o), a) = env.feedbackZero (o, a) := by rw [Unique.eq_default h] rfl @@ -211,7 +212,7 @@ lemma fst_stepKernel (alg : Algorithm π“ž 𝓐 𝓨) (env : Environment π“ž rw [stepKernel, Kernel.fst_compProd] lemma stepKernel_zero (alg : Algorithm π“ž 𝓐 𝓨) (env : Environment π“ž 𝓐 𝓨) (h : Hist π“ž 𝓐 𝓨 0) : - stepKernel alg env 0 h = env.obs0 βŠ—β‚˜ (alg.p0 βŠ—β‚– env.Ξ½0) := by + stepKernel alg env 0 h = env.obsZero βŠ—β‚˜ (alg.policyZero βŠ—β‚– env.feedbackZero) := by rw [Unique.eq_default h, stepKernel, Kernel.compProd_apply_eq_compProd_sectR] congr 1 ext o s hs @@ -475,30 +476,30 @@ lemma hasLaw_history_zero (O : β„• β†’ Ξ© β†’ π“ž) (A : β„• β†’ Ξ© β†’ 𝓐) (Y map_eq := by rw [history_zero, Measure.map_const, measure_univ, one_smul] lemma IsAlgEnvSeqUntil.hasLaw_obs_zero (h : IsAlgEnvSeqUntil O A Y alg env P N) (hN : 0 < N) : - HasLaw (O 0) env.obs0 P := by + HasLaw (O 0) env.obsZero P := by have h0 := h.hasCondDistrib_obs 0 hN rw [history_zero] at h0 exact h0.hasLaw_of_const' lemma IsAlgEnvSeq.hasLaw_obs_zero (h : IsAlgEnvSeq O A Y alg env P) : - HasLaw (O 0) env.obs0 P := + HasLaw (O 0) env.obsZero P := (h.isAlgEnvSeqUntil 1).hasLaw_obs_zero zero_lt_one omit [IsProbabilityMeasure P] in lemma IsAlgEnvSeqUntil.hasCondDistrib_action_zero (h : IsAlgEnvSeqUntil O A Y alg env P N) (hN : 0 < N) : - HasCondDistrib (A 0) (O 0) alg.p0 P := + HasCondDistrib (A 0) (O 0) alg.policyZero P := hasCondDistrib_prodMk_left_unique_iff.mp (h.hasCondDistrib_action 0 hN) omit [IsProbabilityMeasure P] in lemma IsAlgEnvSeq.hasCondDistrib_action_zero (h : IsAlgEnvSeq O A Y alg env P) : - HasCondDistrib (A 0) (O 0) alg.p0 P := + HasCondDistrib (A 0) (O 0) alg.policyZero P := (h.isAlgEnvSeqUntil 1).hasCondDistrib_action_zero zero_lt_one omit [IsProbabilityMeasure P] in lemma IsAlgEnvSeqUntil.hasCondDistrib_feedback_zero (h : IsAlgEnvSeqUntil O A Y alg env P N) (hN : 0 < N) : - HasCondDistrib (Y 0) (fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) env.Ξ½0 P := by + HasCondDistrib (Y 0) (fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) env.feedbackZero P := by have h0 := h.hasCondDistrib_feedback 0 hN rw [history_zero] at h0 exact h0.of_measurableEmbedding_comp_right @@ -506,18 +507,18 @@ lemma IsAlgEnvSeqUntil.hasCondDistrib_feedback_zero (h : IsAlgEnvSeqUntil O A Y omit [IsProbabilityMeasure P] in lemma IsAlgEnvSeq.hasCondDistrib_feedback_zero (h : IsAlgEnvSeq O A Y alg env P) : - HasCondDistrib (Y 0) (fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) env.Ξ½0 P := + HasCondDistrib (Y 0) (fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) env.feedbackZero P := (h.isAlgEnvSeqUntil 1).hasCondDistrib_feedback_zero zero_lt_one lemma IsAlgEnvSeqUntil.hasLaw_step_zero (h : IsAlgEnvSeqUntil O A Y alg env P N) (hN : 0 < N) : - HasLaw (step O A Y 0) (env.obs0 βŠ—β‚˜ (alg.p0 βŠ—β‚– env.Ξ½0)) P := by + HasLaw (step O A Y 0) (env.obsZero βŠ—β‚˜ (alg.policyZero βŠ—β‚– env.feedbackZero)) P := by have h0 := h.hasCondDistrib_step 0 hN rw [history_zero] at h0 rw [← stepKernel_zero alg env default] exact h0.hasLaw_of_const' lemma IsAlgEnvSeq.hasLaw_step_zero (h : IsAlgEnvSeq O A Y alg env P) : - HasLaw (step O A Y 0) (env.obs0 βŠ—β‚˜ (alg.p0 βŠ—β‚– env.Ξ½0)) P := + HasLaw (step O A Y 0) (env.obsZero βŠ—β‚˜ (alg.policyZero βŠ—β‚– env.feedbackZero)) P := (h.isAlgEnvSeqUntil 1).hasLaw_step_zero zero_lt_one end Zero @@ -539,11 +540,11 @@ lemma IsAlgEnvSeq.hasLaw_step_comp (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : HasLaw (step O A Y n) (stepKernel alg env n βˆ˜β‚˜ (P.map (history O A Y n))) P := HasCondDistrib.hasLaw_comp (h.hasCondDistrib_step n) -/-- Conditionally on the event `(O 0, A 0) = p`, the first feedback has law `env.Ξ½0 p`. -/ +/-- Conditionally on the event `(O 0, A 0) = p`, the first feedback has law `env.feedbackZero p`. -/ lemma IsAlgEnvSeq.hasLaw_feedback_zero_cond [MeasurableSingletonClass π“ž] [MeasurableSingletonClass 𝓐] (h : IsAlgEnvSeq O A Y alg env P) {p : π“ž Γ— 𝓐} (hP : P ((fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) ⁻¹' {p}) β‰  0) : - HasLaw (Y 0) (env.Ξ½0 p) P[|(fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) ⁻¹' {p}] := + HasLaw (Y 0) (env.feedbackZero p) P[|(fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) ⁻¹' {p}] := h.hasCondDistrib_feedback_zero.hasLaw_cond (h.measurable_feedback 0) (measurableSet_singleton p) (fun a ha ↦ by rw [Set.mem_singleton_iff.1 ha]) hP diff --git a/LeanMachineLearning/SequentialLearning/AlgorithmDensity.lean b/LeanMachineLearning/SequentialLearning/AlgorithmDensity.lean index 53dd65a8..0175947c 100644 --- a/LeanMachineLearning/SequentialLearning/AlgorithmDensity.lean +++ b/LeanMachineLearning/SequentialLearning/AlgorithmDensity.lean @@ -58,8 +58,8 @@ structure AbsolutelyContinuous (alg algβ‚€ : Algorithm π“ž 𝓐 𝓨) : Prop wh @[inherit_doc AbsolutelyContinuous] scoped notation:50 alg " β‰ͺₐ " algβ‚€ => AbsolutelyContinuous alg algβ‚€ -lemma AbsolutelyContinuous.p0 {alg algβ‚€ : Algorithm π“ž 𝓐 𝓨} (h : alg β‰ͺₐ algβ‚€) (o : π“ž) : - alg.p0 o β‰ͺ algβ‚€.p0 o := +lemma AbsolutelyContinuous.policyZero {alg algβ‚€ : Algorithm π“ž 𝓐 𝓨} (h : alg β‰ͺₐ algβ‚€) (o : π“ž) : + alg.policyZero o β‰ͺ algβ‚€.policyZero o := h.policy 0 (default, o) /-- If the algorithm `alg` is absolutely continuous with respect to the algorithm `algβ‚€` and they diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean b/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean index 00fd20e9..3a612140 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/Markov.lean @@ -74,10 +74,10 @@ lemma policy_apply_eq_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMark alg.policy n (h, o) = alg.policyCondObs n o := by rw [policy_eq_prodMkLeft_policyCondObs, Kernel.prodMkLeft_apply] -lemma p0_eq_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] : - alg.p0 = alg.policyCondObs 0 := by +lemma policyZero_eq_policyCondObs (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsMarkov] : + alg.policyZero = alg.policyCondObs 0 := by ext o : 1 - rw [p0_apply, policy_apply_eq_policyCondObs] + rw [policyZero_apply, policy_apply_eq_policyCondObs] /-- The kernel `alg.policyCondObs n` is determined by the policy at time `n`. The assumption `Nonempty 𝓨` ensures that there are histories of every length. -/ @@ -195,9 +195,9 @@ lemma policy_markov (n : β„•) : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policy n = (ΞΊ n).prodMkLeft (Hist π“ž 𝓐 𝓨 n) := rfl @[simp] -lemma p0_markov : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).p0 = ΞΊ 0 := by +lemma policyZero_markov : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).policyZero = ΞΊ 0 := by ext o : 1 - rw [p0_apply, policy_markov, Kernel.prodMkLeft_apply] + rw [policyZero_apply, policy_markov, Kernel.prodMkLeft_apply] instance : (markov ΞΊ : Algorithm π“ž 𝓐 𝓨).IsMarkov where exists_policy_eq_prodMkLeft := ⟨κ, inferInstance, fun _ ↦ rfl⟩ diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean index 4348030e..26b287bf 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Basic.lean @@ -17,7 +17,7 @@ measure at each iteration. ## Main definitions -* `randomSampling`: The random sampling algorithm that samples from a fixed distribution at +* `Algorithm.const`: The random sampling algorithm that samples from a fixed distribution at each iteration. ## Main statements @@ -43,41 +43,41 @@ open Set in /-- The _Random Sampling_ algorithm, which samples from a fixed probability measure at each iteration. -/ @[simps] -noncomputable def randomSampling (ΞΌ : Measure 𝓐) [IsProbabilityMeasure ΞΌ] : +noncomputable def Algorithm.const (ΞΌ : Measure 𝓐) [IsProbabilityMeasure ΞΌ] : Algorithm π“ž 𝓐 𝓨 where policy _ := Kernel.const _ ΞΌ /-- The random sampling algorithm is the Markov algorithm with the constant kernel `ΞΌ` at every time. -/ -lemma randomSampling_eq_markov : - (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨) = Algorithm.markov (fun _ ↦ Kernel.const π“ž ΞΌ) := rfl +lemma Algorithm.const_eq_markov : + (Algorithm.const ΞΌ : Algorithm π“ž 𝓐 𝓨) = Algorithm.markov (fun _ ↦ Kernel.const π“ž ΞΌ) := rfl -instance : (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).IsMarkov := +instance : (Algorithm.const ΞΌ : Algorithm π“ž 𝓐 𝓨).IsMarkov := inferInstanceAs (Algorithm.markov (fun _ ↦ Kernel.const π“ž ΞΌ) : Algorithm π“ž 𝓐 𝓨).IsMarkov @[simp] -lemma policyCondObs_randomSampling [Nonempty 𝓨] (n : β„•) : - (randomSampling ΞΌ : Algorithm π“ž 𝓐 𝓨).policyCondObs n = Kernel.const π“ž ΞΌ := +lemma policyCondObs_const [Nonempty 𝓨] (n : β„•) : + (Algorithm.const ΞΌ : Algorithm π“ž 𝓐 𝓨).policyCondObs n = Kernel.const π“ž ΞΌ := Algorithm.policyCondObs_eq_of_policy_eq _ rfl -namespace randomSampling +namespace Algorithm.const variable {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} {env : Environment π“ž 𝓐 𝓨} /-- Each action of the random sampling algorithm follows the distribution ΞΌ. -/ -lemma hasLaw_action (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) env P) (n : β„•) : +lemma hasLaw_action (h : IsAlgEnvSeq O A Y (Algorithm.const ΞΌ) env P) (n : β„•) : HasLaw (A n) ΞΌ P := (h.hasCondDistrib_action n).hasLaw_of_const /-- Actions of the random sampling algorithm are mutually independent. -/ -lemma iIndep_action (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) env P) : +lemma iIndep_action (h : IsAlgEnvSeq O A Y (Algorithm.const ΞΌ) env P) : iIndepFun A P := by have hO := h.measurable_obs have hA := h.measurable_action rw [iIndepFun_nat_iff_forall_indepFun (by fun_prop)] intro n have map_eq := (h.hasCondDistrib_action (n + 1)).map_eq - simp only [randomSampling_policy, Measure.compProd_const] at map_eq + simp only [Algorithm.const_policy, Measure.compProd_const] at map_eq have law_eq : P.map (A (n + 1)) = ΞΌ := (hasLaw_action h (n + 1)).map_eq rw [← law_eq, ← indepFun_iff_map_prod_eq_prod_map_map] at map_eq Β· change A (n + 1) βŸ‚α΅’[P] (fun (p : Hist π“ž 𝓐 𝓨 (n + 1) Γ— π“ž) (i : Iic n) ↦ @@ -87,6 +87,6 @@ lemma iIndep_action (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) env P) : Β· exact ((h.measurable_history (n + 1)).prodMk (h.measurable_obs (n + 1))).aemeasurable Β· exact (h.measurable_action (n + 1)).aemeasurable -end randomSampling +end Algorithm.const end Learning diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Tendsto.lean b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Tendsto.lean index 91829c5e..b82d306c 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Tendsto.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/RandomSampling/Tendsto.lean @@ -17,8 +17,8 @@ import LeanMachineLearning.ForMathlib.Topology.Instances.ENNReal.Lemmas /-! # Random Sampling convergence lemmas -This file contains several convergence lemmas for the `randomSampling` algorithm along with an -`evalEnv` environment, that evaluates the actions using a measurable function. +This file contains several convergence lemmas for the `Algorithm.const` algorithm along with an +`Environment.eval` environment, that evaluates the actions using a measurable function. ## Main statements @@ -38,7 +38,7 @@ open Learning MeasureTheory ProbabilityTheory Filter Finset ENNReal open scoped Topology -namespace Learning.randomSampling +namespace Learning.Algorithm.const variable {𝓐 𝓨 Ξ© : Type*} {m𝓐 : MeasurableSpace 𝓐} {m𝓨 : MeasurableSpace 𝓨} {mΞ© : MeasurableSpace Ξ©} {ΞΌ : Measure 𝓐} [IsProbabilityMeasure ΞΌ] {P : Measure Ξ©} @@ -51,18 +51,18 @@ section rewards variable [StandardBorelSpace 𝓨] [Nonempty 𝓨] /-- Each reward follows the distribution ΞΌ.map f. -/ -lemma hasLaw_feeback (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) (evalEnv f hf) P) (n : β„•) : +lemma hasLaw_feeback (h : IsAlgEnvSeq O A Y (Algorithm.const ΞΌ) (Environment.eval f hf) P) (n : β„•) : HasLaw (Y n) (ΞΌ.map f) P := by - refine HasLaw.congr ?_ (feedback_evalEnv_ae_eq_eval_action h n) + refine HasLaw.congr ?_ (feedback_eval_ae_eq_eval_action h n) have hA := h.measurable_action n refine ⟨by fun_prop, ?_⟩ rw [← Measure.map_map hf hA, (hasLaw_action h n).map_eq] /-- Rewards are mutually independent. -/ -lemma iIndep_feedback (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) (evalEnv f hf) P) : +lemma iIndep_feedback (h : IsAlgEnvSeq O A Y (Algorithm.const ΞΌ) (Environment.eval f hf) P) : iIndepFun Y P := have (n : β„•) : f ∘ A n =ᡐ[P] Y n := - (feedback_evalEnv_ae_eq_eval_action h n).symm + (feedback_eval_ae_eq_eval_action h n).symm iIndepFun.congr this <| (iIndep_action h).comp _ (fun _ ↦ hf) end rewards @@ -71,17 +71,17 @@ variable [PseudoMetricSpace 𝓐] [SecondCountableTopology 𝓐] [OpensMeasurabl [ΞΌ.IsOpenPosMeasure] /-- The minimum distance from sampled actions to any point tends to zero. -/ -theorem action_tendsto_any (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) (evalEnv f hf) P) (a : 𝓐) - {Ξ΅ : ℝ} (hΞ΅ : 0 < Ξ΅) : +theorem action_tendsto_any (h : IsAlgEnvSeq O A Y (Algorithm.const ΞΌ) (Environment.eval f hf) P) + (a : 𝓐) {Ξ΅ : ℝ} (hΞ΅ : 0 < Ξ΅) : Tendsto (fun i => P {x | Ξ΅ ≀ (fun (j : Iic i) ↦ dist (A j.1 x) a).min}) atTop (𝓝 0) := by - set randomSampling_alg := randomSampling (π“ž := Unit) (𝓨 := 𝓨) ΞΌ + set const_alg := Algorithm.const (π“ž := Unit) (𝓨 := 𝓨) ΞΌ refine tendsto_zero_of_le (g := fun n ↦ P (β‹‚ i ∈ Iic n, {x | Ξ΅ ≀ dist (A i x) a})) ?_ ?_ Β· have inter_prod (n : β„•) : P (β‹‚ j ∈ Iic n, {x | Ξ΅ ≀ dist (A j x) a}) = ∏ j ∈ Iic n, P {x | Ξ΅ ≀ dist (A j x) a} := by refine iIndepSet.meas_biInter ?_ _ rw [iIndepSet_iff_meas_biInter fun i ↦ ?_] Β· intro s - have iIndep_actions := randomSampling.iIndep_action h + have iIndep_actions := Algorithm.const.iIndep_action h rw [iIndepFun_iff_measure_inter_preimage_eq_mul] at iIndep_actions have meas_dist : βˆ€ i ∈ s, MeasurableSet {x | Ξ΅ ≀ dist x a} := by intro i hs @@ -94,7 +94,7 @@ theorem action_tendsto_any (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) (evalEnv f have prod_law (n : β„•) : ∏ j ∈ Iic n, P {x | Ξ΅ ≀ dist (A j x) a} = ∏ j ∈ Iic n, ΞΌ {x | Ξ΅ ≀ dist x a} := by refine prod_congr rfl fun j hj ↦ ?_ - have hlaw (n : β„•) : HasLaw (A n) ΞΌ P := randomSampling.hasLaw_action h n + have hlaw (n : β„•) : HasLaw (A n) ΞΌ P := Algorithm.const.hasLaw_action h n rw [← (hlaw j).map_eq, P.map_apply] Β· simp Β· exact h.measurable_action j @@ -118,7 +118,7 @@ variable [PseudoMetricSpace 𝓨] [BorelSpace 𝓨] (hfc : Continuous f) /-- The minimum distance from image of actions to any function value tends to zero. -/ lemma image_action_tendsto_any - (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) (evalEnv f hfc.measurable) P) + (h : IsAlgEnvSeq O A Y (Algorithm.const ΞΌ) (Environment.eval f hfc.measurable) P) (a : 𝓐) {Ξ΅ : ℝ} (hΞ΅ : 0 < Ξ΅) : Tendsto (fun i => P {x | Ξ΅ ≀ (fun (j : Iic i) ↦ dist (f (A j.1 x)) (f a)).min}) atTop (𝓝 0) := by @@ -140,20 +140,21 @@ lemma image_action_tendsto_any variable [StandardBorelSpace 𝓨] [Nonempty 𝓨] /-- The minimum distance from rewards to any function value tends to zero. -/ -lemma feedback_tendsto_any (h : IsAlgEnvSeq O A Y (randomSampling ΞΌ) (evalEnv f hfc.measurable) P) +lemma feedback_tendsto_any + (h : IsAlgEnvSeq O A Y (Algorithm.const ΞΌ) (Environment.eval f hfc.measurable) P) (a : 𝓐) {Ξ΅ : ℝ} (hΞ΅ : 0 < Ξ΅) : Tendsto (fun i => P {x | Ξ΅ ≀ (fun (j : Iic i) ↦ dist (Y j.1 x) (f a)).min}) atTop (𝓝 0) := by convert image_action_tendsto_any hfc h a hΞ΅ using 2 with n refine measure_congr ?_ let g : ((Iic n) β†’ 𝓨) β†’ ℝ := fun r ↦ (fun i ↦ dist (r i) (f a)).min - filter_upwards [feedback_evalEnv_ae_eq_eval_action_comp h g] with Ο‰ hΟ‰ + filter_upwards [feedback_eval_ae_eq_eval_action_comp h g] with Ο‰ hΟ‰ simp only [eq_iff_iff] simp [g, hΟ‰] variable {R : β„• β†’ Ξ© β†’ ℝ} {f : 𝓐 β†’ ℝ} (hfc : Continuous f) {a : 𝓐} /-- The minimum image action converges to the function's global minimum. -/ -lemma tendsto_minβ‚€ (h : IsAlgEnvSeq O A R (randomSampling ΞΌ) (evalEnv f hfc.measurable) P) +lemma tendsto_minβ‚€ (h : IsAlgEnvSeq O A R (Algorithm.const ΞΌ) (Environment.eval f hfc.measurable) P) (hf_min : βˆ€ x, f a ≀ f x) : TendstoInMeasure P (fun n Ο‰ ↦ (fun (i : Iic n) ↦ f (A i.1 Ο‰)).min) atTop (fun _ ↦ f a) := by rw [tendstoInMeasure_iff_dist] @@ -174,15 +175,15 @@ lemma tendsto_minβ‚€ (h : IsAlgEnvSeq O A R (randomSampling ΞΌ) (evalEnv f hfc.m grind /-- The minimum reward converges to the function's global minimum. -/ -lemma tendsto_min (h : IsAlgEnvSeq O A R (randomSampling ΞΌ) (evalEnv f hfc.measurable) P) +lemma tendsto_min (h : IsAlgEnvSeq O A R (Algorithm.const ΞΌ) (Environment.eval f hfc.measurable) P) (hf_min : βˆ€ x, f a ≀ f x) : TendstoInMeasure P (fun n Ο‰ ↦ (fun (i : Iic n) ↦ R i.1 Ο‰).min) atTop (fun _ ↦ f a) := by refine TendstoInMeasure.congr_left (fun n ↦ ?_) <| tendsto_minβ‚€ hfc h hf_min - filter_upwards [feedback_evalEnv_ae_eq_eval_action_comp h Function.min] with Ο‰ hΟ‰ + filter_upwards [feedback_eval_ae_eq_eval_action_comp h Function.min] with Ο‰ hΟ‰ rw [← hΟ‰] /-- The maximum image action converges to the function's global maximum. -/ -lemma tendsto_maxβ‚€ (h : IsAlgEnvSeq O A R (randomSampling ΞΌ) (evalEnv f hfc.measurable) P) +lemma tendsto_maxβ‚€ (h : IsAlgEnvSeq O A R (Algorithm.const ΞΌ) (Environment.eval f hfc.measurable) P) (hf_max : βˆ€ x, f x ≀ f a) : TendstoInMeasure P (fun n Ο‰ ↦ (fun (i : Iic n) ↦ f (A i.1 Ο‰)).max) atTop (fun _ ↦ f a) := by rw [tendstoInMeasure_iff_dist] @@ -204,11 +205,11 @@ lemma tendsto_maxβ‚€ (h : IsAlgEnvSeq O A R (randomSampling ΞΌ) (evalEnv f hfc.m grind /-- The maximum reward converges to the function's global maximum. -/ -lemma tendsto_max (h : IsAlgEnvSeq O A R (randomSampling ΞΌ) (evalEnv f hfc.measurable) P) +lemma tendsto_max (h : IsAlgEnvSeq O A R (Algorithm.const ΞΌ) (Environment.eval f hfc.measurable) P) (hf_max : βˆ€ x, f x ≀ f a) : TendstoInMeasure P (fun n Ο‰ ↦ (fun (i : Iic n) ↦ R i.1 Ο‰).max) atTop (fun _ ↦ f a) := by refine TendstoInMeasure.congr_left (fun n ↦ ?_) <| tendsto_maxβ‚€ hfc h hf_max - filter_upwards [feedback_evalEnv_ae_eq_eval_action_comp h Function.max] with Ο‰ hΟ‰ + filter_upwards [feedback_eval_ae_eq_eval_action_comp h Function.max] with Ο‰ hΟ‰ rw [← hΟ‰] -end Learning.randomSampling +end Learning.Algorithm.const diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean b/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean index 9c48b76c..f73d6dde 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/RoundRobin.lean @@ -84,7 +84,7 @@ variable (K) in /-- The Round-Robin algorithm: deterministic algorithm that chooses action `n % K` at time `n`. -/ noncomputable def roundRobinAlgorithm [NeZero K] : Algorithm π“ž (Fin K) 𝓨 := - detAlgorithm (fun n _ ↦ RoundRobin.nextAction K n) (by fun_prop) + Algorithm.deterministic (fun n _ ↦ RoundRobin.nextAction K n) (by fun_prop) end AlgorithmDefinition @@ -99,7 +99,7 @@ variable [NeZero K] {Ξ½ : Kernel (Fin K) 𝓨} [IsMarkovKernel Ξ½] lemma action_ae_eq (n : β„•) (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P (n + 1)) : A n =ᡐ[P] fun _ ↦ nextAction K n := - h.action_detAlgorithm_ae_eq n.lt_succ_self + h.action_deterministic_ae_eq n.lt_succ_self lemma action_zero (h : IsAlgEnvSeqUntil O A Y (roundRobinAlgorithm K) (Environment.bandit Ξ½) P 1) : diff --git a/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean b/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean index 0ae9e703..b636cf58 100644 --- a/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean +++ b/LeanMachineLearning/SequentialLearning/Algorithms/Uniform.lean @@ -38,10 +38,10 @@ variable {π“ž 𝓐 𝓨 : Type*} {mπ“ž : MeasurableSpace π“ž} {m𝓐 : Measur /-- The Uniform algorithm: actions are chosen uniformly at random. -/ noncomputable def uniformAlgorithm [Finite 𝓐] [Nonempty 𝓐] : Algorithm π“ž 𝓐 𝓨 := - randomSampling (uniformOn Set.univ) + Algorithm.const (uniformOn Set.univ) instance [Finite 𝓐] [Nonempty 𝓐] : (uniformAlgorithm : Algorithm π“ž 𝓐 𝓨).IsMarkov := - inferInstanceAs (randomSampling (uniformOn Set.univ) : Algorithm π“ž 𝓐 𝓨).IsMarkov + inferInstanceAs (Algorithm.const (uniformOn Set.univ) : Algorithm π“ž 𝓐 𝓨).IsMarkov lemma absolutelyContinuous_uniformAlgorithm [Finite 𝓐] [Nonempty 𝓐] {alg : Algorithm π“ž 𝓐 𝓨} : alg β‰ͺₐ uniformAlgorithm where diff --git a/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean b/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean index 673b0f21..e177abbd 100644 --- a/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean +++ b/LeanMachineLearning/SequentialLearning/BayesStationaryEnv.lean @@ -94,10 +94,10 @@ lemma feedback_bayesStationaryEnv (n : β„•) : (bayesStationaryEnv Q ΞΊ).feedback n = ΞΊ.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := rfl @[simp] -lemma obs0_bayesStationaryEnv : (bayesStationaryEnv Q ΞΊ).obs0 = Q := rfl +lemma obsZero_bayesStationaryEnv : (bayesStationaryEnv Q ΞΊ).obsZero = Q := rfl @[simp] -lemma Ξ½0_bayesStationaryEnv : (bayesStationaryEnv Q ΞΊ).Ξ½0 = ΞΊ := rfl +lemma feedbackZero_bayesStationaryEnv : (bayesStationaryEnv Q ΞΊ).feedbackZero = ΞΊ := rfl end BayesEnv @@ -184,16 +184,16 @@ lemma hasCondDistrib_feedback' (h : IsBayesAlgEnvSeq Q ΞΊ alg E A Y P) (n : β„•) (h.hasCondDistrib_feedback n).comp_right lemma hasLaw_action_zero (h : IsBayesAlgEnvSeq Q ΞΊ alg E A Y P) : - HasLaw (A 0) (alg.p0 ()) P := by + HasLaw (A 0) (alg.policyZero ()) P := by have h0 : HasCondDistrib (A 0) (fun _ : Ξ© ↦ ((default : Hist Unit 𝓐 𝓨 0), ())) (alg.policy 0) P := by have h1 := h.hasCondDistrib_action' 0 rwa [history_zero] at h1 exact h0.hasLaw_of_const' -/-- The first action is independent of the parameter `E`, and has law `alg.p0 ()`. -/ +/-- The first action is independent of the parameter `E`, and has law `alg.policyZero ()`. -/ lemma hasCondDistrib_action_zero (h : IsBayesAlgEnvSeq Q ΞΊ alg E A Y P) : - HasCondDistrib (A 0) E (Kernel.const _ (alg.p0 ())) P := + HasCondDistrib (A 0) E (Kernel.const _ (alg.policyZero ())) P := hasCondDistrib_prodMk_right_unique_iff.mp (h.hasCondDistrib_action 0) /-- The posterior over the parameter given the empty history is the prior. -/ diff --git a/LeanMachineLearning/SequentialLearning/Comap.lean b/LeanMachineLearning/SequentialLearning/Comap.lean index 78392f5a..bda9378f 100644 --- a/LeanMachineLearning/SequentialLearning/Comap.lean +++ b/LeanMachineLearning/SequentialLearning/Comap.lean @@ -136,12 +136,13 @@ lemma Algorithm.policy_comap (alg : Algorithm π“ž 𝓐 𝓨) (alg.comap F hF).policy n = (alg.policy n).comap (F n) (hF n) := rfl @[simp] -lemma Algorithm.p0_comap (alg : Algorithm π“ž 𝓐 𝓨) +lemma Algorithm.policyZero_comap (alg : Algorithm π“ž 𝓐 𝓨) {F : (n : β„•) β†’ Hist π“ž' 𝓐 𝓨' n Γ— π“ž' β†’ Hist π“ž 𝓐 𝓨 n Γ— π“ž} (hF : βˆ€ n, Measurable (F n)) : - (alg.comap F hF).p0 - = alg.p0.comap (fun o ↦ (F 0 (default, o)).2) (((hF 0).comp measurable_prodMk_left).snd) := by + (alg.comap F hF).policyZero + = alg.policyZero.comap (fun o ↦ (F 0 (default, o)).2) + (((hF 0).comp measurable_prodMk_left).snd) := by ext o : 1 - rw [p0_apply, policy_comap, Kernel.comap_apply, alg.policy_zero, Kernel.comap_apply] + rw [policyZero_apply, policy_comap, Kernel.comap_apply, alg.policy_zero, Kernel.comap_apply] @[simp] lemma Algorithm.comap_id (alg : Algorithm π“ž 𝓐 𝓨) : @@ -172,10 +173,10 @@ lemma Algorithm.policy_comapObs (alg : Algorithm π“ž 𝓐 𝓨) (hf : Measurabl = (alg.policy n).comap (fun p ↦ (Hist.mapObs f p.1, f p.2)) (by fun_prop) := rfl @[simp] -lemma Algorithm.p0_comapObs (alg : Algorithm π“ž 𝓐 𝓨) (hf : Measurable f) : - (alg.comapObs f hf).p0 = alg.p0.comap f hf := by +lemma Algorithm.policyZero_comapObs (alg : Algorithm π“ž 𝓐 𝓨) (hf : Measurable f) : + (alg.comapObs f hf).policyZero = alg.policyZero.comap f hf := by ext o : 1 - rw [p0_apply, policy_comapObs, Kernel.comap_apply, alg.policy_zero, Kernel.comap_apply] + rw [policyZero_apply, policy_comapObs, Kernel.comap_apply, alg.policy_zero, Kernel.comap_apply] @[simp] lemma Algorithm.comapObs_id (alg : Algorithm π“ž 𝓐 𝓨) : alg.comapObs id measurable_id = alg := rfl @@ -205,10 +206,10 @@ lemma Algorithm.policy_comapFeedback (alg : Algorithm π“ž 𝓐 𝓨) (hg : Meas = (alg.policy n).comap (fun p ↦ (Hist.mapFeedback g p.1, p.2)) (by fun_prop) := rfl @[simp] -lemma Algorithm.p0_comapFeedback (alg : Algorithm π“ž 𝓐 𝓨) (hg : Measurable g) : - (alg.comapFeedback g hg).p0 = alg.p0 := by +lemma Algorithm.policyZero_comapFeedback (alg : Algorithm π“ž 𝓐 𝓨) (hg : Measurable g) : + (alg.comapFeedback g hg).policyZero = alg.policyZero := by ext o : 1 - rw [p0_apply, policy_comapFeedback, Kernel.comap_apply, alg.policy_zero, p0_apply] + rw [policyZero_apply, policy_comapFeedback, Kernel.comap_apply, alg.policy_zero, policyZero_apply] @[simp] lemma Algorithm.comapFeedback_id (alg : Algorithm π“ž 𝓐 𝓨) : @@ -255,17 +256,18 @@ lemma Environment.feedback_comap (env : Environment π“ž 𝓐 𝓨) (hF : βˆ€ n, = (env.feedback n).comap (fun p ↦ ((F n p.1.1, p.1.2), f p.2)) (by fun_prop) := rfl @[simp] -lemma Environment.obs0_comap (env : Environment π“ž 𝓐 𝓨) (hF : βˆ€ n, Measurable (F n)) +lemma Environment.obsZero_comap (env : Environment π“ž 𝓐 𝓨) (hF : βˆ€ n, Measurable (F n)) (hf : Measurable f) : - (env.comap F hF f hf).obs0 = env.obs0 := by - rw [Environment.obs0_def, obs_comap, Kernel.comap_apply, env.obs_zero] + (env.comap F hF f hf).obsZero = env.obsZero := by + rw [Environment.obsZero_def, obs_comap, Kernel.comap_apply, env.obs_zero] @[simp] -lemma Environment.Ξ½0_comap (env : Environment π“ž 𝓐 𝓨) (hF : βˆ€ n, Measurable (F n)) +lemma Environment.feedbackZero_comap (env : Environment π“ž 𝓐 𝓨) (hF : βˆ€ n, Measurable (F n)) (hf : Measurable f) : - (env.comap F hF f hf).Ξ½0 = env.Ξ½0.comap (fun p ↦ (p.1, f p.2)) (by fun_prop) := by + (env.comap F hF f hf).feedbackZero + = env.feedbackZero.comap (fun p ↦ (p.1, f p.2)) (by fun_prop) := by ext p : 1 - rw [Environment.Ξ½0_apply, feedback_comap, Kernel.comap_apply, env.feedback_zero, + rw [Environment.feedbackZero_apply, feedback_comap, Kernel.comap_apply, env.feedback_zero, Kernel.comap_apply] @[simp] @@ -299,15 +301,16 @@ lemma Environment.feedback_comapAction (env : Environment π“ž 𝓐 𝓨) (hf : (fun p ↦ ((Hist.mapAction f p.1.1, p.1.2), f p.2)) (by fun_prop) := rfl @[simp] -lemma Environment.obs0_comapAction (env : Environment π“ž 𝓐 𝓨) (hf : Measurable f) : - (env.comapAction f hf).obs0 = env.obs0 := by - rw [Environment.obs0_def, obs_comapAction, Kernel.comap_apply, env.obs_zero] +lemma Environment.obsZero_comapAction (env : Environment π“ž 𝓐 𝓨) (hf : Measurable f) : + (env.comapAction f hf).obsZero = env.obsZero := by + rw [Environment.obsZero_def, obs_comapAction, Kernel.comap_apply, env.obs_zero] @[simp] -lemma Environment.Ξ½0_comapAction (env : Environment π“ž 𝓐 𝓨) (hf : Measurable f) : - (env.comapAction f hf).Ξ½0 = env.Ξ½0.comap (fun p ↦ (p.1, f p.2)) (by fun_prop) := by +lemma Environment.feedbackZero_comapAction (env : Environment π“ž 𝓐 𝓨) (hf : Measurable f) : + (env.comapAction f hf).feedbackZero + = env.feedbackZero.comap (fun p ↦ (p.1, f p.2)) (by fun_prop) := by ext p : 1 - rw [Environment.Ξ½0_apply, feedback_comapAction, Kernel.comap_apply, env.feedback_zero, + rw [Environment.feedbackZero_apply, feedback_comapAction, Kernel.comap_apply, env.feedback_zero, Kernel.comap_apply] @[simp] @@ -341,11 +344,11 @@ lemma Algorithm.policy_congr (alg : Algorithm π“ž 𝓐 𝓨) (eπ“ž : π“ž ≃ (fun p ↦ (Hist.map eπ“ž.symm e𝓐.symm e𝓨.symm p.1, eπ“ž.symm p.2)) (by fun_prop) := rfl @[simp] -lemma Algorithm.p0_congr (alg : Algorithm π“ž 𝓐 𝓨) (eπ“ž : π“ž ≃ᡐ π“ž') (e𝓐 : 𝓐 ≃ᡐ 𝓐') +lemma Algorithm.policyZero_congr (alg : Algorithm π“ž 𝓐 𝓨) (eπ“ž : π“ž ≃ᡐ π“ž') (e𝓐 : 𝓐 ≃ᡐ 𝓐') (e𝓨 : 𝓨 ≃ᡐ 𝓨') : - (alg.congr eπ“ž e𝓐 e𝓨).p0 = (alg.p0.map e𝓐).comap eπ“ž.symm eπ“ž.symm.measurable := by + (alg.congr eπ“ž e𝓐 e𝓨).policyZero = (alg.policyZero.map e𝓐).comap eπ“ž.symm eπ“ž.symm.measurable := by ext o : 1 - rw [p0_apply, policy_congr, Kernel.comap_apply, Kernel.map_apply _ e𝓐.measurable, + rw [policyZero_apply, policy_congr, Kernel.comap_apply, Kernel.map_apply _ e𝓐.measurable, alg.policy_zero, Kernel.comap_apply, Kernel.map_apply _ e𝓐.measurable] @[simp] @@ -402,19 +405,19 @@ lemma Environment.feedback_congr (env : Environment π“ž 𝓐 𝓨) (eπ“ž : (by fun_prop) := rfl @[simp] -lemma Environment.obs0_congr (env : Environment π“ž 𝓐 𝓨) (eπ“ž : π“ž ≃ᡐ π“ž') (e𝓐 : 𝓐 ≃ᡐ 𝓐') +lemma Environment.obsZero_congr (env : Environment π“ž 𝓐 𝓨) (eπ“ž : π“ž ≃ᡐ π“ž') (e𝓐 : 𝓐 ≃ᡐ 𝓐') (e𝓨 : 𝓨 ≃ᡐ 𝓨') : - (env.congr eπ“ž e𝓐 e𝓨).obs0 = env.obs0.map eπ“ž := by - rw [Environment.obs0_def, obs_congr, Kernel.comap_apply, Kernel.map_apply _ eπ“ž.measurable, + (env.congr eπ“ž e𝓐 e𝓨).obsZero = env.obsZero.map eπ“ž := by + rw [Environment.obsZero_def, obs_congr, Kernel.comap_apply, Kernel.map_apply _ eπ“ž.measurable, env.obs_zero] @[simp] -lemma Environment.Ξ½0_congr (env : Environment π“ž 𝓐 𝓨) (eπ“ž : π“ž ≃ᡐ π“ž') (e𝓐 : 𝓐 ≃ᡐ 𝓐') +lemma Environment.feedbackZero_congr (env : Environment π“ž 𝓐 𝓨) (eπ“ž : π“ž ≃ᡐ π“ž') (e𝓐 : 𝓐 ≃ᡐ 𝓐') (e𝓨 : 𝓨 ≃ᡐ 𝓨') : - (env.congr eπ“ž e𝓐 e𝓨).Ξ½0 - = (env.Ξ½0.map e𝓨).comap (fun p ↦ (eπ“ž.symm p.1, e𝓐.symm p.2)) (by fun_prop) := by + (env.congr eπ“ž e𝓐 e𝓨).feedbackZero + = (env.feedbackZero.map e𝓨).comap (fun p ↦ (eπ“ž.symm p.1, e𝓐.symm p.2)) (by fun_prop) := by ext p : 1 - rw [Environment.Ξ½0_apply, feedback_congr, Kernel.comap_apply, + rw [Environment.feedbackZero_apply, feedback_congr, Kernel.comap_apply, Kernel.map_apply _ e𝓨.measurable, env.feedback_zero, Kernel.comap_apply, Kernel.map_apply _ e𝓨.measurable] diff --git a/LeanMachineLearning/SequentialLearning/Deterministic.lean b/LeanMachineLearning/SequentialLearning/Deterministic.lean index 66784d36..ab7824b0 100644 --- a/LeanMachineLearning/SequentialLearning/Deterministic.lean +++ b/LeanMachineLearning/SequentialLearning/Deterministic.lean @@ -17,28 +17,33 @@ kernel. Similarly, a deterministic environment gives feedback in a deterministic ## Main definitions -We introduce two typeclasses `IsDeterministicAlg` and `IsDeterministicEnv` to express that -an algorithm or an environment is deterministic. We also give definitions for the initial action +We introduce two typeclasses `Algorithm.IsDeterministic` and +`Environment.HasDeterministicFeedback` to express that an algorithm is deterministic or that an +environment gives deterministic feedback. We also give definitions for the initial action and the next action of a deterministic algorithm, and for the feedback functions of a deterministic environment. Finally, we give a construction of a deterministic algorithm and environment from measurable functions. -* `IsDeterministicAlg alg`: a typeclass expressing that the algorithm `alg` is deterministic. -* `IsDeterministicEnv env`: a typeclass expressing that the environment `env` is deterministic. -* `nextAction alg n`: the function that gives the action of a deterministic algorithm `alg` - at step `n`, as a function of the history before `n` and of the observation at step `n`. -* `actionZero alg`: the initial action of a deterministic algorithm `alg`, as a function of the - first observation. This is `nextAction alg 0` applied to the empty history. -* `feedbackFun env n`: the function that gives the feedback of a deterministic environment `env` - at step `n`, as a function of the history, the current observation and the current action. -* `feedbackFunZero env`: the function that gives the initial feedback of a deterministic - environment `env`. This is `feedbackFun env 0` applied to the empty history. - -* `detAlgorithm nextA h_next`: a deterministic algorithm that chooses its action +* `Algorithm.IsDeterministic alg`: a typeclass expressing that the algorithm `alg` is + deterministic. +* `Environment.HasDeterministicFeedback env`: a typeclass expressing that the feedback of the + environment `env` is deterministic. +* `Algorithm.nextAction alg n`: the function that gives the action of a deterministic algorithm + `alg` at step `n`, as a function of the history before `n` and of the observation at step `n`. +* `Algorithm.actionZero alg`: the initial action of a deterministic algorithm `alg`, as a function + of the first observation. This is `alg.nextAction 0` applied to the empty history. +* `Environment.feedbackFun env n`: the function that gives the feedback of a deterministic + environment `env` at step `n`, as a function of the history, the current observation and the + current action. +* `Environment.feedbackFunZero env`: the function that gives the initial feedback of a + deterministic environment `env`. This is `env.feedbackFun 0` applied to the empty history. + +* `Algorithm.deterministic nextA h_next`: a deterministic algorithm that chooses its action according to the measurable function `nextA` (with proof of measurability `h_next`). The initial action is `fun o ↦ nextA 0 (default, o)`. -* `detEnvironment obs f hf`: a deterministic environment with observation kernels `obs`, that gives - feedback according to the measurable function `f` (with proof of measurability `hf`). +* `Environment.detFeedback obs f hf`: an environment with observation kernels `obs`, that gives + deterministic feedback according to the measurable function `f` (with proof of measurability + `hf`). -/ @@ -55,178 +60,186 @@ variable {π“ž 𝓐 𝓨 : Type*} {mπ“ž : MeasurableSpace π“ž} {m𝓐 : Measur /-- An algorithm is deterministic if its actions are determined by measurable functions of the history and of the current observation (and not possibly random kernels). -/ -class IsDeterministicAlg (alg : Algorithm π“ž 𝓐 𝓨) : Prop where +class Algorithm.IsDeterministic (alg : Algorithm π“ž 𝓐 𝓨) : Prop where exists_nextAction n : βˆƒ (nextAction : (Hist π“ž 𝓐 𝓨 n Γ— π“ž) β†’ 𝓐) (h_meas : Measurable nextAction), alg.policy n = Kernel.deterministic nextAction h_meas +namespace Algorithm + /-- The action of a deterministic algorithm at step `n`, as a function of the history before `n` and of the observation at step `n`. -/ noncomputable -def nextAction (alg : Algorithm π“ž 𝓐 𝓨) [h_det : IsDeterministicAlg alg] (n : β„•) : +def nextAction (alg : Algorithm π“ž 𝓐 𝓨) [h_det : alg.IsDeterministic] (n : β„•) : (Hist π“ž 𝓐 𝓨 n Γ— π“ž) β†’ 𝓐 := (h_det.exists_nextAction n).choose /-- The initial action of a deterministic algorithm, as a function of the first observation. -/ noncomputable -def actionZero (alg : Algorithm π“ž 𝓐 𝓨) [IsDeterministicAlg alg] : π“ž β†’ 𝓐 := - fun o ↦ nextAction alg 0 (default, o) +def actionZero (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsDeterministic] : π“ž β†’ 𝓐 := + fun o ↦ alg.nextAction 0 (default, o) @[fun_prop] -lemma measurable_nextAction (alg : Algorithm π“ž 𝓐 𝓨) [IsDeterministicAlg alg] (n : β„•) : - Measurable (nextAction alg n) := - (IsDeterministicAlg.exists_nextAction n).choose_spec.choose +lemma measurable_nextAction (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsDeterministic] (n : β„•) : + Measurable (alg.nextAction n) := + (IsDeterministic.exists_nextAction n).choose_spec.choose @[fun_prop] -lemma measurable_actionZero (alg : Algorithm π“ž 𝓐 𝓨) [IsDeterministicAlg alg] : - Measurable (actionZero alg) := - (measurable_nextAction alg 0).comp (measurable_const.prodMk measurable_id) +lemma measurable_actionZero (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsDeterministic] : + Measurable alg.actionZero := + (alg.measurable_nextAction 0).comp (measurable_const.prodMk measurable_id) -lemma policy_eq_deterministic (alg : Algorithm π“ž 𝓐 𝓨) [h_det : IsDeterministicAlg alg] (n : β„•) : - alg.policy n = Kernel.deterministic (nextAction alg n) (measurable_nextAction alg n) := - (IsDeterministicAlg.exists_nextAction n).choose_spec.choose_spec +lemma policy_eq_deterministic (alg : Algorithm π“ž 𝓐 𝓨) [h_det : alg.IsDeterministic] (n : β„•) : + alg.policy n = Kernel.deterministic (alg.nextAction n) (alg.measurable_nextAction n) := + (IsDeterministic.exists_nextAction n).choose_spec.choose_spec -lemma nextAction_zero (alg : Algorithm π“ž 𝓐 𝓨) [IsDeterministicAlg alg] (h : Hist π“ž 𝓐 𝓨 0) +lemma nextAction_zero (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsDeterministic] (h : Hist π“ž 𝓐 𝓨 0) (o : π“ž) : - nextAction alg 0 (h, o) = actionZero alg o := by + alg.nextAction 0 (h, o) = alg.actionZero o := by rw [Unique.eq_default h] rfl -lemma p0_eq_deterministic (alg : Algorithm π“ž 𝓐 𝓨) [IsDeterministicAlg alg] : - alg.p0 = Kernel.deterministic (actionZero alg) (measurable_actionZero alg) := by +lemma policyZero_eq_deterministic (alg : Algorithm π“ž 𝓐 𝓨) [alg.IsDeterministic] : + alg.policyZero = Kernel.deterministic alg.actionZero alg.measurable_actionZero := by ext o : 1 - rw [Algorithm.p0_apply, policy_eq_deterministic, Kernel.deterministic_apply, + rw [policyZero_apply, policy_eq_deterministic, Kernel.deterministic_apply, Kernel.deterministic_apply] rfl -namespace IsDeterministicAlg +end Algorithm + +namespace Algorithm.IsDeterministic variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm π“ž 𝓐 𝓨} {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} [IsFiniteMeasure P] {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} {n N : β„•} lemma action_ae_eq_of_IsAlgEnvSeqUntil [MeasurableEq 𝓐] - [h_det : IsDeterministicAlg alg] (h : IsAlgEnvSeqUntil O A Y alg env P N) (hn : n < N) : - A n =ᡐ[P] fun Ο‰ ↦ nextAction alg n (history O A Y n Ο‰, O n Ο‰) := by + [h_det : alg.IsDeterministic] (h : IsAlgEnvSeqUntil O A Y alg env P N) (hn : n < N) : + A n =ᡐ[P] fun Ο‰ ↦ alg.nextAction n (history O A Y n Ο‰, O n Ο‰) := by have h_eq := (h.hasCondDistrib_action n hn) - rw [policy_eq_deterministic alg n] at h_eq + rw [alg.policy_eq_deterministic n] at h_eq have hO := h.measurable_obs have hA := h.measurable_action have hY := h.measurable_feedback exact ae_eq_of_hasCondDistrib_deterministic (measurable_nextAction _ _) (by fun_prop) (by fun_prop) h_eq -lemma action_zero_of_IsAlgEnvSeqUntil [MeasurableEq 𝓐] [h_det : IsDeterministicAlg alg] +lemma action_zero_of_IsAlgEnvSeqUntil [MeasurableEq 𝓐] [h_det : alg.IsDeterministic] (h : IsAlgEnvSeqUntil O A Y alg env P N) (hN : 0 < N) : - A 0 =ᡐ[P] fun Ο‰ ↦ actionZero alg (O 0 Ο‰) := by + A 0 =ᡐ[P] fun Ο‰ ↦ alg.actionZero (O 0 Ο‰) := by filter_upwards [action_ae_eq_of_IsAlgEnvSeqUntil h hN] with Ο‰ hΟ‰ rw [hΟ‰, nextAction_zero] -lemma hasCondDistrib_action_zero_of_IsAlgEnvSeqUntil [h_det : IsDeterministicAlg alg] +lemma hasCondDistrib_action_zero_of_IsAlgEnvSeqUntil [h_det : alg.IsDeterministic] (h : IsAlgEnvSeqUntil O A Y alg env P N) (hN : 0 < N) : HasCondDistrib (A 0) (O 0) - (Kernel.deterministic (actionZero alg) (measurable_actionZero alg)) P := by - rw [← p0_eq_deterministic] + (Kernel.deterministic alg.actionZero alg.measurable_actionZero) P := by + rw [← policyZero_eq_deterministic] exact h.hasCondDistrib_action_zero hN -lemma hasCondDistrib_action_zero [h_det : IsDeterministicAlg alg] +lemma hasCondDistrib_action_zero [h_det : alg.IsDeterministic] (h : IsAlgEnvSeq O A Y alg env P) : HasCondDistrib (A 0) (O 0) - (Kernel.deterministic (actionZero alg) (measurable_actionZero alg)) P := + (Kernel.deterministic alg.actionZero alg.measurable_actionZero) P := hasCondDistrib_action_zero_of_IsAlgEnvSeqUntil (h.isAlgEnvSeqUntil 1) zero_lt_one -lemma action_ae_eq [MeasurableEq 𝓐] [h_det : IsDeterministicAlg alg] +lemma action_ae_eq [MeasurableEq 𝓐] [h_det : alg.IsDeterministic] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : - A n =ᡐ[P] fun Ο‰ ↦ nextAction alg n (history O A Y n Ο‰, O n Ο‰) := + A n =ᡐ[P] fun Ο‰ ↦ alg.nextAction n (history O A Y n Ο‰, O n Ο‰) := action_ae_eq_of_IsAlgEnvSeqUntil (h.isAlgEnvSeqUntil (n + 1)) n.lt_succ_self -lemma action_zero_ae_eq [MeasurableEq 𝓐] [h_det : IsDeterministicAlg alg] +lemma action_zero_ae_eq [MeasurableEq 𝓐] [h_det : alg.IsDeterministic] (h : IsAlgEnvSeq O A Y alg env P) : - A 0 =ᡐ[P] fun Ο‰ ↦ actionZero alg (O 0 Ο‰) := + A 0 =ᡐ[P] fun Ο‰ ↦ alg.actionZero (O 0 Ο‰) := action_zero_of_IsAlgEnvSeqUntil (h.isAlgEnvSeqUntil 1) zero_lt_one -lemma action_ae_all_eq [MeasurableEq 𝓐] [h_det : IsDeterministicAlg alg] +lemma action_ae_all_eq [MeasurableEq 𝓐] [h_det : alg.IsDeterministic] (h : IsAlgEnvSeq O A Y alg env P) : - βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, A n Ο‰ = nextAction alg n (history O A Y n Ο‰, O n Ο‰) := + βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, A n Ο‰ = alg.nextAction n (history O A Y n Ο‰, O n Ο‰) := ae_all_iff.mpr (action_ae_eq h) -end IsDeterministicAlg +end Algorithm.IsDeterministic -/-- An environment is deterministic if its feedbacks are determined by measurable functions of -the history, the observation and the action (and not possibly random kernels). -/ -class IsDeterministicEnv (env : Environment π“ž 𝓐 𝓨) : Prop where +/-- An environment has deterministic feedback if its feedbacks are determined by measurable +functions of the history, the observation and the action (and not possibly random kernels). -/ +class Environment.HasDeterministicFeedback (env : Environment π“ž 𝓐 𝓨) : Prop where exists_f : βˆ€ n, βˆƒ (f : ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐) β†’ 𝓨) (hf : Measurable f), env.feedback n = Kernel.deterministic f hf +namespace Environment + /-- The feedback function of a deterministic environment at step `n`. -/ noncomputable -def feedbackFun (env : Environment π“ž 𝓐 𝓨) [h_det : IsDeterministicEnv env] (n : β„•) : +def feedbackFun (env : Environment π“ž 𝓐 𝓨) [h_det : env.HasDeterministicFeedback] (n : β„•) : ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐) β†’ 𝓨 := (h_det.exists_f n).choose @[fun_prop] -lemma measurable_feedbackFun (env : Environment π“ž 𝓐 𝓨) [IsDeterministicEnv env] (n : β„•) : - Measurable (feedbackFun env n) := - (IsDeterministicEnv.exists_f n).choose_spec.choose +lemma measurable_feedbackFun (env : Environment π“ž 𝓐 𝓨) [env.HasDeterministicFeedback] (n : β„•) : + Measurable (env.feedbackFun n) := + (HasDeterministicFeedback.exists_f n).choose_spec.choose -lemma feedback_eq_deterministic (env : Environment π“ž 𝓐 𝓨) [IsDeterministicEnv env] (n : β„•) : - env.feedback n = Kernel.deterministic (feedbackFun env n) (measurable_feedbackFun env n) := - (IsDeterministicEnv.exists_f n).choose_spec.choose_spec +lemma feedback_eq_deterministic (env : Environment π“ž 𝓐 𝓨) [env.HasDeterministicFeedback] (n : β„•) : + env.feedback n = Kernel.deterministic (env.feedbackFun n) (env.measurable_feedbackFun n) := + (HasDeterministicFeedback.exists_f n).choose_spec.choose_spec /-- The initial feedback function of a deterministic environment, as a function of the first observation and the first action. -/ noncomputable -def feedbackFunZero (env : Environment π“ž 𝓐 𝓨) [IsDeterministicEnv env] : π“ž Γ— 𝓐 β†’ 𝓨 := - fun p ↦ feedbackFun env 0 ((default, p.1), p.2) +def feedbackFunZero (env : Environment π“ž 𝓐 𝓨) [env.HasDeterministicFeedback] : π“ž Γ— 𝓐 β†’ 𝓨 := + fun p ↦ env.feedbackFun 0 ((default, p.1), p.2) @[fun_prop] -lemma measurable_feedbackFunZero (env : Environment π“ž 𝓐 𝓨) [IsDeterministicEnv env] : - Measurable (feedbackFunZero env) := - (measurable_feedbackFun env 0).comp +lemma measurable_feedbackFunZero (env : Environment π“ž 𝓐 𝓨) [env.HasDeterministicFeedback] : + Measurable env.feedbackFunZero := + (env.measurable_feedbackFun 0).comp ((measurable_const.prodMk measurable_fst).prodMk measurable_snd) -lemma feedbackFun_zero (env : Environment π“ž 𝓐 𝓨) [IsDeterministicEnv env] (h : Hist π“ž 𝓐 𝓨 0) +lemma feedbackFun_zero (env : Environment π“ž 𝓐 𝓨) [env.HasDeterministicFeedback] (h : Hist π“ž 𝓐 𝓨 0) (o : π“ž) (a : 𝓐) : - feedbackFun env 0 ((h, o), a) = feedbackFunZero env (o, a) := by + env.feedbackFun 0 ((h, o), a) = env.feedbackFunZero (o, a) := by rw [Unique.eq_default h] rfl -lemma Ξ½0_eq_deterministic (env : Environment π“ž 𝓐 𝓨) [IsDeterministicEnv env] : - env.Ξ½0 = Kernel.deterministic (feedbackFunZero env) (measurable_feedbackFunZero env) := by +lemma feedbackZero_eq_deterministic (env : Environment π“ž 𝓐 𝓨) [env.HasDeterministicFeedback] : + env.feedbackZero = Kernel.deterministic env.feedbackFunZero env.measurable_feedbackFunZero := by ext p : 1 - rw [Environment.Ξ½0_def, Kernel.comap_apply, feedback_eq_deterministic, Kernel.deterministic_apply, + rw [feedbackZero_def, Kernel.comap_apply, feedback_eq_deterministic, Kernel.deterministic_apply, Kernel.deterministic_apply] rfl -namespace IsDeterministicEnv +end Environment + +namespace Environment.HasDeterministicFeedback variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm π“ž 𝓐 𝓨} {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} [IsFiniteMeasure P] {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} -lemma hasCondDistrib_feedback [h_det : IsDeterministicEnv env] +lemma hasCondDistrib_feedback [h_det : env.HasDeterministicFeedback] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : HasCondDistrib (Y n) (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) - (Kernel.deterministic (feedbackFun env n) (measurable_feedbackFun env n)) P := by + (Kernel.deterministic (env.feedbackFun n) (env.measurable_feedbackFun n)) P := by rw [← feedback_eq_deterministic] exact h.hasCondDistrib_feedback n -lemma hasCondDistrib_feedback_zero [h_det : IsDeterministicEnv env] +lemma hasCondDistrib_feedback_zero [h_det : env.HasDeterministicFeedback] (h : IsAlgEnvSeq O A Y alg env P) : HasCondDistrib (Y 0) (fun Ο‰ ↦ (O 0 Ο‰, A 0 Ο‰)) - (Kernel.deterministic (feedbackFunZero env) (measurable_feedbackFunZero env)) P := by - rw [← Ξ½0_eq_deterministic] + (Kernel.deterministic env.feedbackFunZero env.measurable_feedbackFunZero) P := by + rw [← feedbackZero_eq_deterministic] exact h.hasCondDistrib_feedback_zero -lemma feedback_ae_eq [MeasurableEq 𝓨] [h_det : IsDeterministicEnv env] +lemma feedback_ae_eq [MeasurableEq 𝓨] [h_det : env.HasDeterministicFeedback] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : - Y n =ᡐ[P] fun Ο‰ ↦ feedbackFun env n ((history O A Y n Ο‰, O n Ο‰), A n Ο‰) := by + Y n =ᡐ[P] fun Ο‰ ↦ env.feedbackFun n ((history O A Y n Ο‰, O n Ο‰), A n Ο‰) := by have hO := h.measurable_obs have hA := h.measurable_action have hY := h.measurable_feedback exact ae_eq_of_hasCondDistrib_deterministic (measurable_feedbackFun _ _) (by fun_prop) (by fun_prop) (hasCondDistrib_feedback h n) -end IsDeterministicEnv +end Environment.HasDeterministicFeedback variable {nextA : (n : β„•) β†’ (Hist π“ž 𝓐 𝓨 n Γ— π“ž) β†’ 𝓐} {h_next : βˆ€ n, Measurable (nextA n)} {env : Environment π“ž 𝓐 𝓨} @@ -237,38 +250,38 @@ variable {nextA : (n : β„•) β†’ (Hist π“ž 𝓐 𝓨 n Γ— π“ž) β†’ 𝓐} {h_nex The initial action is `fun o ↦ nextA 0 (default, o)`. -/ @[simps] noncomputable -def detAlgorithm (nextA : (n : β„•) β†’ (Hist π“ž 𝓐 𝓨 n Γ— π“ž) β†’ 𝓐) +def Algorithm.deterministic (nextA : (n : β„•) β†’ (Hist π“ž 𝓐 𝓨 n Γ— π“ž) β†’ 𝓐) (h_next : βˆ€ n, Measurable (nextA n)) : Algorithm π“ž 𝓐 𝓨 where policy n := Kernel.deterministic (nextA n) (h_next n) -instance : IsDeterministicAlg (detAlgorithm nextA h_next) where +instance : (Algorithm.deterministic nextA h_next).IsDeterministic where exists_nextAction n := ⟨nextA n, h_next n, rfl⟩ @[simp] -lemma p0_detAlgorithm : - (detAlgorithm nextA h_next).p0 +lemma policyZero_deterministic : + (Algorithm.deterministic nextA h_next).policyZero = Kernel.deterministic (fun o ↦ nextA 0 (default, o)) ((h_next 0).comp (measurable_const.prodMk measurable_id)) := by ext o : 1 - rw [Algorithm.p0_apply, detAlgorithm_policy, Kernel.deterministic_apply, + rw [Algorithm.policyZero_apply, Algorithm.deterministic_policy, Kernel.deterministic_apply, Kernel.deterministic_apply] @[simp] -lemma nextAction_detAlgorithm [MeasurableSpace.SeparatesPoints 𝓐] (n : β„•) : - nextAction (detAlgorithm nextA h_next) n = nextA n := by - have h_eq := policy_eq_deterministic (detAlgorithm nextA h_next) n - simpa [detAlgorithm] using h_eq.symm +lemma nextAction_deterministic [MeasurableSpace.SeparatesPoints 𝓐] (n : β„•) : + (Algorithm.deterministic nextA h_next).nextAction n = nextA n := by + have h_eq := (Algorithm.deterministic nextA h_next).policy_eq_deterministic n + simpa [Algorithm.deterministic] using h_eq.symm @[simp] -lemma actionZero_detAlgorithm [MeasurableSpace.SeparatesPoints 𝓐] : - actionZero (detAlgorithm nextA h_next) = fun o ↦ nextA 0 (default, o) := by - unfold actionZero - rw [nextAction_detAlgorithm] +lemma actionZero_deterministic [MeasurableSpace.SeparatesPoints 𝓐] : + (Algorithm.deterministic nextA h_next).actionZero = fun o ↦ nextA 0 (default, o) := by + unfold Algorithm.actionZero + rw [nextAction_deterministic] /-- A deterministic environment, where the feedback is given by evaluating fixed measurable functions. -/ -noncomputable def detEnvironment (obs : (n : β„•) β†’ Kernel (Hist π“ž 𝓐 𝓨 n) π“ž) +noncomputable def Environment.detFeedback (obs : (n : β„•) β†’ Kernel (Hist π“ž 𝓐 𝓨 n) π“ž) [βˆ€ n, IsMarkovKernel (obs n)] (f : (n : β„•) β†’ ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐) β†’ 𝓨) (hf : βˆ€ n, Measurable (f n)) : Environment π“ž 𝓐 𝓨 where @@ -276,25 +289,26 @@ noncomputable def detEnvironment (obs : (n : β„•) β†’ Kernel (Hist π“ž 𝓐 feedback n := (Kernel.deterministic (f n) (hf n)) @[simp] -lemma obs_detEnvironment (n : β„•) : (detEnvironment obs f hf).obs n = obs n := rfl +lemma obs_detFeedback (n : β„•) : (Environment.detFeedback obs f hf).obs n = obs n := rfl @[simp] -lemma feedback_detEnvironment (n : β„•) : - (detEnvironment obs f hf).feedback n = Kernel.deterministic (f n) (hf n) := rfl +lemma feedback_detFeedback (n : β„•) : + (Environment.detFeedback obs f hf).feedback n = Kernel.deterministic (f n) (hf n) := rfl -instance : IsDeterministicEnv (detEnvironment obs f hf) where +instance : (Environment.detFeedback obs f hf).HasDeterministicFeedback where exists_f n := ⟨f n, hf n, rfl⟩ @[simp] -lemma feedbackFun_detEnvironment [MeasurableSpace.SeparatesPoints 𝓨] (n : β„•) : - feedbackFun (detEnvironment obs f hf) n = f n := by - simpa [detEnvironment] using (feedback_eq_deterministic (detEnvironment obs f hf) n).symm +lemma feedbackFun_detFeedback [MeasurableSpace.SeparatesPoints 𝓨] (n : β„•) : + (Environment.detFeedback obs f hf).feedbackFun n = f n := by + simpa [Environment.detFeedback] using + ((Environment.detFeedback obs f hf).feedback_eq_deterministic n).symm @[simp] -lemma feedbackFunZero_detEnvironment [MeasurableSpace.SeparatesPoints 𝓨] : - feedbackFunZero (detEnvironment obs f hf) = fun p ↦ f 0 ((default, p.1), p.2) := by - unfold feedbackFunZero - rw [feedbackFun_detEnvironment] +lemma feedbackFunZero_detFeedback [MeasurableSpace.SeparatesPoints 𝓨] : + (Environment.detFeedback obs f hf).feedbackFunZero = fun p ↦ f 0 ((default, p.1), p.2) := by + unfold Environment.feedbackFunZero + rw [feedbackFun_detFeedback] namespace IsAlgEnvSeq @@ -303,28 +317,28 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {P : Measure Ξ©} [IsProbabilityMeasure P] {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} -lemma hasCondDistrib_action_zero_detAlgorithm - (h : IsAlgEnvSeq O A Y (detAlgorithm nextA h_next) env P) : +lemma hasCondDistrib_action_zero_deterministic + (h : IsAlgEnvSeq O A Y (Algorithm.deterministic nextA h_next) env P) : HasCondDistrib (A 0) (O 0) (Kernel.deterministic (fun o ↦ nextA 0 (default, o)) ((h_next 0).comp (measurable_const.prodMk measurable_id))) P := by - rw [← p0_detAlgorithm] + rw [← policyZero_deterministic] exact h.hasCondDistrib_action_zero -lemma action_detAlgorithm_ae_eq [MeasurableEq 𝓐] - (h : IsAlgEnvSeq O A Y (detAlgorithm nextA h_next) env P) (n : β„•) : +lemma action_deterministic_ae_eq [MeasurableEq 𝓐] + (h : IsAlgEnvSeq O A Y (Algorithm.deterministic nextA h_next) env P) (n : β„•) : A n =ᡐ[P] fun Ο‰ ↦ nextA n (history O A Y n Ο‰, O n Ο‰) := - (IsDeterministicAlg.action_ae_eq h n).trans (by simp) + (Algorithm.IsDeterministic.action_ae_eq h n).trans (by simp) -lemma action_zero_detAlgorithm [MeasurableEq 𝓐] - (h : IsAlgEnvSeq O A Y (detAlgorithm nextA h_next) env P) : +lemma action_zero_deterministic [MeasurableEq 𝓐] + (h : IsAlgEnvSeq O A Y (Algorithm.deterministic nextA h_next) env P) : A 0 =ᡐ[P] fun Ο‰ ↦ nextA 0 (default, O 0 Ο‰) := - (IsDeterministicAlg.action_zero_ae_eq h).trans (by simp) + (Algorithm.IsDeterministic.action_zero_ae_eq h).trans (by simp) -lemma action_detAlgorithm_ae_all_eq [MeasurableEq 𝓐] - (h : IsAlgEnvSeq O A Y (detAlgorithm nextA h_next) env P) : +lemma action_deterministic_ae_all_eq [MeasurableEq 𝓐] + (h : IsAlgEnvSeq O A Y (Algorithm.deterministic nextA h_next) env P) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, A n Ο‰ = nextA n (history O A Y n Ο‰, O n Ο‰) := - ae_all_iff.mpr (action_detAlgorithm_ae_eq h) + ae_all_iff.mpr (action_deterministic_ae_eq h) end IsAlgEnvSeq @@ -335,23 +349,23 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {P : Measure Ξ©} [IsProbabilityMeasure P] {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} {N n : β„•} -lemma hasCondDistrib_action_zero_detAlgorithm - (h : IsAlgEnvSeqUntil O A Y (detAlgorithm nextA h_next) env P N) (hN : 0 < N) : +lemma hasCondDistrib_action_zero_deterministic + (h : IsAlgEnvSeqUntil O A Y (Algorithm.deterministic nextA h_next) env P N) (hN : 0 < N) : HasCondDistrib (A 0) (O 0) (Kernel.deterministic (fun o ↦ nextA 0 (default, o)) ((h_next 0).comp (measurable_const.prodMk measurable_id))) P := by - rw [← p0_detAlgorithm] + rw [← policyZero_deterministic] exact h.hasCondDistrib_action_zero hN -lemma action_detAlgorithm_ae_eq [MeasurableEq 𝓐] - (h : IsAlgEnvSeqUntil O A Y (detAlgorithm nextA h_next) env P N) (hn : n < N) : +lemma action_deterministic_ae_eq [MeasurableEq 𝓐] + (h : IsAlgEnvSeqUntil O A Y (Algorithm.deterministic nextA h_next) env P N) (hn : n < N) : A n =ᡐ[P] fun Ο‰ ↦ nextA n (history O A Y n Ο‰, O n Ο‰) := - (IsDeterministicAlg.action_ae_eq_of_IsAlgEnvSeqUntil h hn).trans (by simp) + (Algorithm.IsDeterministic.action_ae_eq_of_IsAlgEnvSeqUntil h hn).trans (by simp) -lemma action_zero_detAlgorithm [MeasurableEq 𝓐] - (h : IsAlgEnvSeqUntil O A Y (detAlgorithm nextA h_next) env P N) (hN : 0 < N) : +lemma action_zero_deterministic [MeasurableEq 𝓐] + (h : IsAlgEnvSeqUntil O A Y (Algorithm.deterministic nextA h_next) env P N) (hN : 0 < N) : A 0 =ᡐ[P] fun Ο‰ ↦ nextA 0 (default, O 0 Ο‰) := - (IsDeterministicAlg.action_zero_of_IsAlgEnvSeqUntil h hN).trans (by simp) + (Algorithm.IsDeterministic.action_zero_of_IsAlgEnvSeqUntil h hN).trans (by simp) end IsAlgEnvSeqUntil diff --git a/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean b/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean index 34477026..3d94838f 100644 --- a/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean +++ b/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean @@ -12,24 +12,25 @@ public import LeanMachineLearning.ForMathlib.Probability.Independence.CondDistri /-! # Function evaluation environments -We define two environments, `onlineEvalEnv` and `evalEnv`, where the feedback is given by evaluating -a measurable function at the chosen action. The first one allows the function to change at every -time step, while the second one uses a fixed function at every time step. +We define two environments, `Environment.evalSeq` and `Environment.eval`, where the feedback is +given by evaluating a measurable function at the chosen action. The first one allows the function +to change at every time step, while the second one uses a fixed function at every time step. ## Main definitions -* `onlineEvalEnv g hg`: A stationary environment where the feedback at time `n` is given by a +* `Environment.evalSeq g hg`: A stationary environment where the feedback at time `n` is given by a deterministic kernel that evaluates the measurable function `g n` at the chosen action. -* `evalEnv f hf`: A stationary environment where the feedback is given by a deterministic kernel - that evaluates a fixed measurable function `f` at the chosen action. +* `Environment.eval f hf`: A stationary environment where the feedback is given by a deterministic + kernel that evaluates a fixed measurable function `f` at the chosen action. -They both satisfy the typeclasses `IsObliviousEnv` and `IsDeterministicEnv`. +They both satisfy the typeclasses `Environment.IsOblivious` and +`Environment.HasDeterministicFeedback`. ## Main statements -* `forall_feedback_onlineEvalEnv_ae_eq_eval_action`: For almost all `Ο‰`, the feedback at time `n` is +* `forall_feedback_evalSeq_ae_eq_eval_action`: For almost all `Ο‰`, the feedback at time `n` is equal to `g n` evaluated at the action taken at time `n`. -* `forall_feedback_evalEnv_ae_eq_eval_action`: For almost all `Ο‰`, the feedback at time `n` is equal +* `forall_feedback_eval_ae_eq_eval_action`: For almost all `Ο‰`, the feedback at time `n` is equal to `f` evaluated at the action taken at time `n`. -/ @@ -46,33 +47,33 @@ variable {𝓐 𝓨 : Type*} {m𝓐 : MeasurableSpace 𝓐} {m𝓨 : MeasurableS /-- The evaluation environment where the feedback is given by evaluating a fixed measurable function `f` at the chosen action. -/ -noncomputable def onlineEvalEnv (g : β„• β†’ 𝓐 β†’ 𝓨) (hg : βˆ€ n, Measurable (g n)) := +noncomputable def Environment.evalSeq (g : β„• β†’ 𝓐 β†’ 𝓨) (hg : βˆ€ n, Measurable (g n)) := Environment.banditSeq (fun n ↦ Kernel.deterministic (g n) (hg n)) -instance : IsObliviousEnv (onlineEvalEnv g hg) := - inferInstanceAs (IsObliviousEnv (Environment.banditSeq fun n ↦ Kernel.deterministic (g n) (hg n))) +instance : (Environment.evalSeq g hg).IsOblivious := + inferInstanceAs (Environment.banditSeq fun n ↦ Kernel.deterministic (g n) (hg n)).IsOblivious -instance : IsDeterministicEnv (onlineEvalEnv g hg) where +instance : (Environment.evalSeq g hg).HasDeterministicFeedback where exists_f n := ⟨fun p ↦ g n p.2, by fun_prop, rfl⟩ @[simp] -lemma feedbackCondObsAction_onlineEvalEnv (n : β„•) : - (onlineEvalEnv g hg).feedbackCondObsAction n +lemma feedbackCondObsAction_evalSeq (n : β„•) : + (Environment.evalSeq g hg).feedbackCondObsAction n = Kernel.deterministic (fun p ↦ g n p.2) (by fun_prop) := by - simp [onlineEvalEnv] + simp [Environment.evalSeq] @[simp] -lemma feedbackFun_onlineEvalEnv [MeasurableSpace.SeparatesPoints 𝓨] (n : β„•) : - feedbackFun (onlineEvalEnv g hg) n = fun p ↦ g n p.2 := by - have h_eq := feedback_eq_deterministic (onlineEvalEnv g hg) n - simpa only [onlineEvalEnv, feedback_banditSeq, Kernel.prodMkLeft_deterministic, +lemma feedbackFun_evalSeq [MeasurableSpace.SeparatesPoints 𝓨] (n : β„•) : + (Environment.evalSeq g hg).feedbackFun n = fun p ↦ g n p.2 := by + have h_eq := (Environment.evalSeq g hg).feedback_eq_deterministic n + simpa only [Environment.evalSeq, feedback_banditSeq, Kernel.prodMkLeft_deterministic, Kernel.deterministic_inj] using h_eq.symm @[simp] -lemma feedbackFunZero_onlineEvalEnv [MeasurableSpace.SeparatesPoints 𝓨] : - feedbackFunZero (onlineEvalEnv g hg) = fun p ↦ g 0 p.2 := by - unfold feedbackFunZero - rw [feedbackFun_onlineEvalEnv] +lemma feedbackFunZero_evalSeq [MeasurableSpace.SeparatesPoints 𝓨] : + (Environment.evalSeq g hg).feedbackFunZero = fun p ↦ g 0 p.2 := by + unfold Environment.feedbackFunZero + rw [feedbackFun_evalSeq] section OnlineEvalEnv @@ -81,48 +82,50 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm Unit 𝓐 𝓨 {P : Measure Ξ©} [IsProbabilityMeasure P] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} -lemma hasCondDistrib_feedback_onlineEvalEnv - (h : IsAlgEnvSeq O A Y alg (onlineEvalEnv g hg) P) (n : β„•) : +lemma hasCondDistrib_feedback_evalSeq + (h : IsAlgEnvSeq O A Y alg (Environment.evalSeq g hg) P) (n : β„•) : HasCondDistrib (Y n) (A n) (Kernel.deterministic (g n) (hg n)) P := h.hasCondDistrib_feedback_banditSeq n -lemma feedback_onlineEvalEnv_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (onlineEvalEnv g hg) P) (n : β„•) : +lemma feedback_evalSeq_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] + (h : IsAlgEnvSeq O A Y alg (Environment.evalSeq g hg) P) (n : β„•) : Y n =ᡐ[P] g n ∘ A n := ae_eq_of_condDistrib_eq_deterministic (hg n) (h.measurable_action n).aemeasurable (h.measurable_feedback n).aemeasurable - (hasCondDistrib_feedback_onlineEvalEnv h n).condDistrib_eq + (hasCondDistrib_feedback_evalSeq h n).condDistrib_eq -lemma forall_feedback_onlineEvalEnv_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (onlineEvalEnv g hg) P) : +lemma forall_feedback_evalSeq_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] + (h : IsAlgEnvSeq O A Y alg (Environment.evalSeq g hg) P) : βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, Y n Ο‰ = g n (A n Ο‰) := by rw [ae_all_iff] intro n - exact feedback_onlineEvalEnv_ae_eq_eval_action h n + exact feedback_evalSeq_ae_eq_eval_action h n end OnlineEvalEnv /-- The evaluation environment where the feedback is given by evaluating a fixed measurable function `f` at the chosen action. -/ -noncomputable def evalEnv (f : 𝓐 β†’ 𝓨) (hf : Measurable f) := onlineEvalEnv (fun _ ↦ f) (fun _ ↦ hf) +noncomputable def Environment.eval (f : 𝓐 β†’ 𝓨) (hf : Measurable f) := + Environment.evalSeq (fun _ ↦ f) (fun _ ↦ hf) -instance : IsObliviousEnv (evalEnv f hf) := by unfold evalEnv; infer_instance +instance : (Environment.eval f hf).IsOblivious := by unfold Environment.eval; infer_instance -instance : IsDeterministicEnv (evalEnv f hf) := by unfold evalEnv; infer_instance +instance : (Environment.eval f hf).HasDeterministicFeedback := by + unfold Environment.eval; infer_instance @[simp] -lemma feedbackCondObsAction_evalEnv (n : β„•) : - (evalEnv f hf).feedbackCondObsAction n +lemma feedbackCondObsAction_eval (n : β„•) : + (Environment.eval f hf).feedbackCondObsAction n = Kernel.deterministic (fun p ↦ f p.2) (by fun_prop) := by - simp [evalEnv] + simp [Environment.eval] @[simp] -lemma feedbackFunZero_evalEnv [MeasurableSpace.SeparatesPoints 𝓨] : - feedbackFunZero (evalEnv f hf) = fun p ↦ f p.2 := by simp [evalEnv] +lemma feedbackFunZero_eval [MeasurableSpace.SeparatesPoints 𝓨] : + (Environment.eval f hf).feedbackFunZero = fun p ↦ f p.2 := by simp [Environment.eval] @[simp] -lemma feedbackFun_evalEnv [MeasurableSpace.SeparatesPoints 𝓨] (n : β„•) : - feedbackFun (evalEnv f hf) n = fun p ↦ f p.2 := by simp [evalEnv] +lemma feedbackFun_eval [MeasurableSpace.SeparatesPoints 𝓨] (n : β„•) : + (Environment.eval f hf).feedbackFun n = fun p ↦ f p.2 := by simp [Environment.eval] section EvalEnv @@ -131,23 +134,23 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm Unit 𝓐 𝓨 {P : Measure Ξ©} [IsProbabilityMeasure P] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} -lemma hasCondDistrib_feedback_evalEnv (h : IsAlgEnvSeq O A Y alg (evalEnv f hf) P) (n : β„•) : +lemma hasCondDistrib_feedback_eval (h : IsAlgEnvSeq O A Y alg (Environment.eval f hf) P) (n : β„•) : HasCondDistrib (Y n) (A n) (Kernel.deterministic f hf) P := h.hasCondDistrib_feedback_banditSeq n -lemma feedback_evalEnv_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (evalEnv f hf) P) (n : β„•) : - Y n =ᡐ[P] f ∘ A n := feedback_onlineEvalEnv_ae_eq_eval_action h n +lemma feedback_eval_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] + (h : IsAlgEnvSeq O A Y alg (Environment.eval f hf) P) (n : β„•) : + Y n =ᡐ[P] f ∘ A n := feedback_evalSeq_ae_eq_eval_action h n -lemma forall_feedback_evalEnv_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (evalEnv f hf) P) : - βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, Y n Ο‰ = f (A n Ο‰) := forall_feedback_onlineEvalEnv_ae_eq_eval_action h +lemma forall_feedback_eval_ae_eq_eval_action [StandardBorelSpace 𝓨] [Nonempty 𝓨] + (h : IsAlgEnvSeq O A Y alg (Environment.eval f hf) P) : + βˆ€α΅ Ο‰ βˆ‚P, βˆ€ n, Y n Ο‰ = f (A n Ο‰) := forall_feedback_evalSeq_ae_eq_eval_action h open Finset in -lemma feedback_evalEnv_ae_eq_eval_action_comp {Ξ² : Type*} [StandardBorelSpace 𝓨] [Nonempty 𝓨] - (h : IsAlgEnvSeq O A Y alg (evalEnv f hf) P) {n : β„•} (g : (Iic n β†’ 𝓨) β†’ Ξ²) : +lemma feedback_eval_ae_eq_eval_action_comp {Ξ² : Type*} [StandardBorelSpace 𝓨] [Nonempty 𝓨] + (h : IsAlgEnvSeq O A Y alg (Environment.eval f hf) P) {n : β„•} (g : (Iic n β†’ 𝓨) β†’ Ξ²) : βˆ€α΅ Ο‰ βˆ‚P, g (fun i ↦ Y i Ο‰) = g (fun i ↦ f (A i Ο‰)) := by - filter_upwards [forall_feedback_evalEnv_ae_eq_eval_action h] with Ο‰ hΟ‰ + filter_upwards [forall_feedback_eval_ae_eq_eval_action h] with Ο‰ hΟ‰ simp_rw [hΟ‰] end EvalEnv diff --git a/LeanMachineLearning/SequentialLearning/IonescuTulceaSpace.lean b/LeanMachineLearning/SequentialLearning/IonescuTulceaSpace.lean index 515c6874..9601a064 100644 --- a/LeanMachineLearning/SequentialLearning/IonescuTulceaSpace.lean +++ b/LeanMachineLearning/SequentialLearning/IonescuTulceaSpace.lean @@ -345,19 +345,20 @@ lemma isAlgEnvSeq_trajMeasure (alg : Algorithm π“ž 𝓐 𝓨) (env : Environmen hasCondDistrib_feedback n := hasCondDistrib_feedback alg env n lemma hasLaw_step_zero (alg : Algorithm π“ž 𝓐 𝓨) (env : Environment π“ž 𝓐 𝓨) : - HasLaw (step 0) (env.obs0 βŠ—β‚˜ (alg.p0 βŠ—β‚– env.Ξ½0)) (trajMeasure alg env) := + HasLaw (step 0) (env.obsZero βŠ—β‚˜ (alg.policyZero βŠ—β‚– env.feedbackZero)) (trajMeasure alg env) := (isAlgEnvSeq_trajMeasure alg env).hasLaw_step_zero lemma hasLaw_obs_zero (alg : Algorithm π“ž 𝓐 𝓨) (env : Environment π“ž 𝓐 𝓨) : - HasLaw (obs 0) env.obs0 (trajMeasure alg env) := + HasLaw (obs 0) env.obsZero (trajMeasure alg env) := (isAlgEnvSeq_trajMeasure alg env).hasLaw_obs_zero lemma hasCondDistrib_action_zero (alg : Algorithm π“ž 𝓐 𝓨) (env : Environment π“ž 𝓐 𝓨) : - HasCondDistrib (action 0) (obs 0) alg.p0 (trajMeasure alg env) := + HasCondDistrib (action 0) (obs 0) alg.policyZero (trajMeasure alg env) := (isAlgEnvSeq_trajMeasure alg env).hasCondDistrib_action_zero lemma hasCondDistrib_feedback_zero (alg : Algorithm π“ž 𝓐 𝓨) (env : Environment π“ž 𝓐 𝓨) : - HasCondDistrib (feedback 0) (fun Ο‰ ↦ (obs 0 Ο‰, action 0 Ο‰)) env.Ξ½0 (trajMeasure alg env) := + HasCondDistrib (feedback 0) (fun Ο‰ ↦ (obs 0 Ο‰, action 0 Ο‰)) env.feedbackZero + (trajMeasure alg env) := (isAlgEnvSeq_trajMeasure alg env).hasCondDistrib_feedback_zero end Laws diff --git a/LeanMachineLearning/SequentialLearning/Means.lean b/LeanMachineLearning/SequentialLearning/Means.lean index 59e0c2d1..dc1418e7 100644 --- a/LeanMachineLearning/SequentialLearning/Means.lean +++ b/LeanMachineLearning/SequentialLearning/Means.lean @@ -65,23 +65,23 @@ noncomputable def Environment.means (env : Environment π“ž 𝓐 𝓨) (O : β„• @[simp] lemma means_zero (env : Environment π“ž 𝓐 𝓨) (O : β„• β†’ Ξ© β†’ π“ž) (A : β„• β†’ Ξ© β†’ 𝓐) (Y : β„• β†’ Ξ© β†’ 𝓨) (k : 𝓐) (Ο‰ : Ξ©) : - env.means O A Y k 0 Ο‰ = (env.Ξ½0 (O 0 Ο‰, k))[id] := by + env.means O A Y k 0 Ο‰ = (env.feedbackZero (O 0 Ο‰, k))[id] := by simp [Environment.means, Environment.measure, Environment.feedback_zero] @[simp] -lemma means_of_isObliviousEnv [IsObliviousEnv env] (O : β„• β†’ Ξ© β†’ π“ž) (A : β„• β†’ Ξ© β†’ 𝓐) +lemma means_of_isOblivious [env.IsOblivious] (O : β„• β†’ Ξ© β†’ π“ž) (A : β„• β†’ Ξ© β†’ 𝓐) (Y : β„• β†’ Ξ© β†’ 𝓨) (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : env.means O A Y k n Ο‰ = (env.feedbackCondObsAction n (O n Ο‰, k))[id] := by simp [Environment.means, Environment.measure, env.feedback_eq_comap_feedbackCondObsAction, Kernel.comap_apply] -lemma means_obliviousEnv (ΞΌ : β„• β†’ Measure π“ž) [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] +lemma means_oblivious (ΞΌ : β„• β†’ Measure π“ž) [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] (Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : - (obliviousEnv ΞΌ Ξ½).means O A Y k n Ο‰ = (Ξ½ n (O n Ο‰, k))[id] := by simp + (Environment.oblivious ΞΌ Ξ½).means O A Y k n Ο‰ = (Ξ½ n (O n Ο‰, k))[id] := by simp -lemma means_stationaryEnv (ΞΌ : Measure π“ž) [IsProbabilityMeasure ΞΌ] (Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨) +lemma means_stationary (ΞΌ : Measure π“ž) [IsProbabilityMeasure ΞΌ] (Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨) [IsMarkovKernel Ξ½] (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : - (stationaryEnv ΞΌ Ξ½).means O A Y k n Ο‰ = (Ξ½ (O n Ο‰, k))[id] := by simp + (Environment.stationary ΞΌ Ξ½).means O A Y k n Ο‰ = (Ξ½ (O n Ο‰, k))[id] := by simp lemma means_banditSeq (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] {O : β„• β†’ Ξ© β†’ Unit} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : diff --git a/LeanMachineLearning/SequentialLearning/README.md b/LeanMachineLearning/SequentialLearning/README.md index d7a0bd64..f21fcca3 100644 --- a/LeanMachineLearning/SequentialLearning/README.md +++ b/LeanMachineLearning/SequentialLearning/README.md @@ -28,10 +28,10 @@ In many applications, some of those kernels are deterministic, or do not depend We detail here the naming conventions for the various constructors, predicates and accessors that are used in the library. Generic constructions live in the `Algorithm` and `Environment` namespaces. -Predicates are root-level `Is…Alg` / `Is…Env` classes when they carry an accessor, and namespaced `Prop` definitions otherwise. +Predicates live in the `Algorithm` and `Environment` namespaces: they are classes when they carry an accessor, and `Prop` definitions otherwise. Accessors are namespaced so that dot notation works. -All time zero accessors are root-level `…0` definitions. Example: `Algorithm.policy0`. +All time zero accessors are `…Zero` definitions. Example: `Algorithm.policyZero`. ## Algorithms @@ -39,13 +39,13 @@ The policy at round `n` can depend on `n`, on the history at `n` and the current Not stochastic: `Algorithm.IsDeterministic`, `Algorithm.deterministic` -No observation: `Algorithm.IgnoresObs`, `Algorithm.comapObs fun _ ↦ ()` +No observation: `Algorithm.IgnoresObs` (MISSING), `Algorithm.comapObs fun _ ↦ ()` No history: `Algorithm.IsMarkov`, `Algorithm.markov` -No history, no observation: `Algorithm.IsOpenLoop`, `Algorithm.openLoop`, `Algorithm.ofSeq` (det version) +No history, no observation: `Algorithm.IsOpenLoop` (MISSING), `Algorithm.openLoop` (MISSING), `Algorithm.ofSeq` (MISSING) (det version) -Not time-dependent, no history: `Algorithm.IsStationary`, `Algorithm.stationary` +Not time-dependent, no history: `Algorithm.IsStationary` (MISSING), `Algorithm.stationary` (MISSING) No time, no history, no observation: `Algorithm.const` @@ -58,15 +58,15 @@ It can be deterministic or stochastic. In general, the dependence on history is the same for both kernels. -All for obs, no action for feedback: `Environment.FeedbackIgnoresAction`, `Environment.adversary`. +All for obs, no action for feedback: `Environment.FeedbackIgnoresAction` (MISSING), `Environment.adversary` (MISSING). No history for obs and feedback: `Environment.IsOblivious`, `Environment.oblivious`. -No time, no history for obs and feedback: `Environment.IsStationary`, `Environment.stationary`. +No time, no history for obs and feedback: `Environment.IsStationary` (MISSING), `Environment.stationary`. -No time, last round of history for obs, not history for feedback: `Environment.IsMarkov`, `Environment.markov`. +No time, last round of history for obs, not history for feedback: `Environment.IsMarkov` (MISSING), `Environment.markov` (MISSING). -Determinism: `Environment.HasDeterministicObs`, `Environment.HasDeterministicFeedback`. +Determinism: `Environment.HasDeterministicObs` (MISSING), `Environment.HasDeterministicFeedback`, `Environment.detFeedback` (any observation kernels, deterministic feedback). ### Obs = Unit @@ -79,9 +79,9 @@ No history: `Environment.banditSeq` and `Environment.bandit` (no time). No history, deterministic: `Environment.evalSeq` and `Environment.eval` (no time). -No history, no action (only time): `Environment.indep` and `Environment.ofSeq` (deterministic). +No history, no action (only time): `Environment.indep` (MISSING) and `Environment.ofSeq` (MISSING) (deterministic). -Nothing: `Environment.const` (stochastic). The deterministic version is probably not useful. +Nothing: `Environment.const` (MISSING) (stochastic). The deterministic version is probably not useful. ## Examples diff --git a/LeanMachineLearning/SequentialLearning/StationaryEnv.lean b/LeanMachineLearning/SequentialLearning/StationaryEnv.lean index 1bc11152..da693d3b 100644 --- a/LeanMachineLearning/SequentialLearning/StationaryEnv.lean +++ b/LeanMachineLearning/SequentialLearning/StationaryEnv.lean @@ -21,11 +21,11 @@ stationary. ## Main definitions -We define a `Prop`-valued typeclass `IsObliviousEnv` to express that an environment is oblivious, -and we define constructors for oblivious environments, with and without observations. +We define a `Prop`-valued typeclass `Environment.IsOblivious` to express that an environment is +oblivious, and we define constructors for oblivious environments, with and without observations. Typeclass and related definitions: -* `IsObliviousEnv env`: the environment `env` is oblivious. +* `Environment.IsOblivious env`: the environment `env` is oblivious. * `Environment.obsLaw env n`: the law of the observation at time `n` in an oblivious environment `env`. * `Environment.feedbackCondObsAction env n`: the kernel representing the conditional distribution @@ -33,11 +33,11 @@ Typeclass and related definitions: environment `env`. Constructors for oblivious environments: -* `obliviousEnv ΞΌ Ξ½`: the oblivious environment in which the observation at time `n` has law `ΞΌ n` - and the feedback at time `n` is drawn from the Markov kernel `Ξ½ n : Kernel (π“ž Γ— 𝓐) 𝓨` applied to - the observation and the action at time `n`. -* `stationaryEnv ΞΌ Ξ½`: the oblivious environment with constant sequences: the observations have - law `ΞΌ` and the feedback is drawn from `Ξ½` applied to the observation and the action. +* `Environment.oblivious ΞΌ Ξ½`: the oblivious environment in which the observation at time `n` has + law `ΞΌ n` and the feedback at time `n` is drawn from the Markov kernel `Ξ½ n : Kernel (π“ž Γ— 𝓐) 𝓨` + applied to the observation and the action at time `n`. +* `Environment.stationary ΞΌ Ξ½`: the oblivious environment with constant sequences: the observations + have law `ΞΌ` and the feedback is drawn from `Ξ½` applied to the observation and the action. * `Environment.banditSeq Ξ½`, `Environment.bandit Ξ½`: the versions without observations (`π“ž = Unit`), in which the feedback at time `n` is drawn from `Ξ½ n : Kernel 𝓐 𝓨` (respectively from `Ξ½ : Kernel 𝓐 𝓨`) applied to the action at time `n`. @@ -58,7 +58,7 @@ variable {π“ž 𝓐 𝓨 : Type*} {mπ“ž : MeasurableSpace π“ž} {m𝓐 : Measur /-- An environment is oblivious if the distributions of the next observation and feedback don't depend on the past history: the observation at time `n` has a fixed law, and the feedback at time `n` depends only on the observation and the action at time `n`. -/ -class IsObliviousEnv (env : Environment π“ž 𝓐 𝓨) : Prop where +class Environment.IsOblivious (env : Environment π“ž 𝓐 𝓨) : Prop where exists_obs_eq_const : βˆƒ ΞΌ : β„• β†’ Measure π“ž, (βˆ€ n, IsProbabilityMeasure (ΞΌ n)) ∧ βˆ€ n, env.obs n = Kernel.const _ (ΞΌ n) exists_feedback_eq_comap : βˆƒ Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨, (βˆ€ n, IsMarkovKernel (Ξ½ n)) ∧ @@ -68,53 +68,53 @@ namespace Environment /-- The law of the observation at time `n` in an oblivious environment. -/ noncomputable -def obsLaw (env : Environment π“ž 𝓐 𝓨) [h_obl : IsObliviousEnv env] (n : β„•) : Measure π“ž := +def obsLaw (env : Environment π“ž 𝓐 𝓨) [h_obl : env.IsOblivious] (n : β„•) : Measure π“ž := h_obl.exists_obs_eq_const.choose n -instance (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] (n : β„•) : +instance (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] (n : β„•) : IsProbabilityMeasure (env.obsLaw n) := - IsObliviousEnv.exists_obs_eq_const.choose_spec.1 n + IsOblivious.exists_obs_eq_const.choose_spec.1 n -lemma obs_eq_const_obsLaw (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] (n : β„•) : +lemma obs_eq_const_obsLaw (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] (n : β„•) : env.obs n = Kernel.const _ (env.obsLaw n) := - IsObliviousEnv.exists_obs_eq_const.choose_spec.2 n + IsOblivious.exists_obs_eq_const.choose_spec.2 n -lemma obs0_eq_obsLaw (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] : - env.obs0 = env.obsLaw 0 := by - rw [Environment.obs0_def, obs_eq_const_obsLaw, Kernel.const_apply] +lemma obsZero_eq_obsLaw (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] : + env.obsZero = env.obsLaw 0 := by + rw [Environment.obsZero_def, obs_eq_const_obsLaw, Kernel.const_apply] /-- The kernel representing the conditional distribution of the feedback given the observation and the action at time `n` in an oblivious environment. -/ noncomputable -def feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [h_obl : IsObliviousEnv env] (n : β„•) : +def feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [h_obl : env.IsOblivious] (n : β„•) : Kernel (π“ž Γ— 𝓐) 𝓨 := h_obl.exists_feedback_eq_comap.choose n -instance (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] (n : β„•) : +instance (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] (n : β„•) : IsMarkovKernel (env.feedbackCondObsAction n) := - IsObliviousEnv.exists_feedback_eq_comap.choose_spec.1 n + IsOblivious.exists_feedback_eq_comap.choose_spec.1 n -lemma feedback_eq_comap_feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] +lemma feedback_eq_comap_feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] (n : β„•) : env.feedback n = (env.feedbackCondObsAction n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := - IsObliviousEnv.exists_feedback_eq_comap.choose_spec.2 n + IsOblivious.exists_feedback_eq_comap.choose_spec.2 n -lemma Ξ½0_eq_feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [IsObliviousEnv env] : - env.Ξ½0 = env.feedbackCondObsAction 0 := by +lemma feedbackZero_eq_feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] : + env.feedbackZero = env.feedbackCondObsAction 0 := by ext p : 1 - rw [Environment.Ξ½0_def, Kernel.comap_apply, feedback_eq_comap_feedbackCondObsAction, + rw [Environment.feedbackZero_def, Kernel.comap_apply, feedback_eq_comap_feedbackCondObsAction, Kernel.comap_apply] end Environment -namespace IsObliviousEnv +namespace Environment.IsOblivious variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm π“ž 𝓐 𝓨} {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} {n N : β„•} /-- The observation at time `n` has law `env.obsLaw n`. -/ -lemma hasLaw_obs [IsProbabilityMeasure P] [IsObliviousEnv env] +lemma hasLaw_obs [IsProbabilityMeasure P] [env.IsOblivious] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : HasLaw (O n) (env.obsLaw n) P := by have h' := h.hasCondDistrib_obs n @@ -123,7 +123,7 @@ lemma hasLaw_obs [IsProbabilityMeasure P] [IsObliviousEnv env] variable [IsFiniteMeasure P] -lemma hasCondDistrib_feedback_history_action [IsObliviousEnv env] +lemma hasCondDistrib_feedback_history_action [env.IsOblivious] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : HasCondDistrib (Y n) (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ((env.feedbackCondObsAction n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) @@ -133,14 +133,14 @@ lemma hasCondDistrib_feedback_history_action [IsObliviousEnv env] /-- The conditional distribution of the feedback at time `n` given the observation and the action at time `n` is `env.feedbackCondObsAction n`. -/ -lemma hasCondDistrib_feedback [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : +lemma hasCondDistrib_feedback [env.IsOblivious] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) (env.feedbackCondObsAction n) P := (hasCondDistrib_feedback_history_action h n).comp_right /-- Conditionally on an event determined by the history before time `n`, the observation and the action at time `n`, on which the observation-action pair is equal to `b`, the feedback at time `n` has law `env.feedbackCondObsAction n b`. -/ -lemma hasLaw_feedback_cond [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) +lemma hasLaw_feedback_cond [env.IsOblivious] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) {s : Set ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐)} (hs : MeasurableSet s) {b : π“ž Γ— 𝓐} (hsb : βˆ€ u ∈ s, (u.1.2, u.2) = b) (hP : P ((fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s) β‰  0) : @@ -153,7 +153,7 @@ lemma hasLaw_feedback_cond [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P /-- Conditionally on an event determined by the history before time `n`, the observation and the action at time `n`, on which the observation-action pair is constant, the feedback at time `n` is independent of the history before time `n`, the observation and the action at time `n`. -/ -lemma indepFun_history_action_feedback_cond [IsObliviousEnv env] +lemma indepFun_history_action_feedback_cond [env.IsOblivious] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) {s : Set ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐)} (hs : MeasurableSet s) {b : π“ž Γ— 𝓐} (hsb : βˆ€ u ∈ s, (u.1.2, u.2) = b) : @@ -172,7 +172,7 @@ variable [StandardBorelSpace π“ž] [Nonempty π“ž] [StandardBorelSpace 𝓐] [No /-- The feedback at time `n` is conditionally independent of the history before time `n`, given the observation and the action at time `n`. -/ lemma condIndepFun_feedback_history [StandardBorelSpace Ξ©] - [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + [env.IsOblivious] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : Y n βŸ‚α΅’[fun Ο‰ ↦ (O n Ο‰, A n Ο‰), (h.measurable_obs n).prodMk (h.measurable_action n); P] history O A Y n := by have hO := h.measurable_obs @@ -192,7 +192,7 @@ lemma condIndepFun_feedback_history [StandardBorelSpace Ξ©] /-- The feedback at time `n` is conditionally independent of the history before time `n`, the observation and the action at time `n`, given the observation and the action at time `n`. -/ lemma condIndepFun_feedback_history_obs_action [StandardBorelSpace Ξ©] - [IsObliviousEnv env] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + [env.IsOblivious] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : Y n βŸ‚α΅’[fun Ο‰ ↦ (O n Ο‰, A n Ο‰), (h.measurable_obs n).prodMk (h.measurable_action n); P] (fun Ο‰ ↦ (history O A Y n Ο‰, (O n Ο‰, A n Ο‰))) := by have hO := h.measurable_obs @@ -200,7 +200,7 @@ lemma condIndepFun_feedback_history_obs_action [StandardBorelSpace Ξ©] have hY := h.measurable_feedback exact (condIndepFun_feedback_history h n).prod_right (by fun_prop) (by fun_prop) (by fun_prop) -end IsObliviousEnv +end Environment.IsOblivious section Oblivious @@ -211,49 +211,49 @@ variable {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] at time `n` is drawn from `Ξ½ n` applied to the observation and the action at time `n`, whatever the past history. -/ noncomputable -def obliviousEnv (ΞΌ : β„• β†’ Measure π“ž) [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] +def Environment.oblivious (ΞΌ : β„• β†’ Measure π“ž) [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] (Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : Environment π“ž 𝓐 𝓨 where obs n := Kernel.const _ (ΞΌ n) feedback n := (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) @[simp] -lemma obs_obliviousEnv (n : β„•) : (obliviousEnv ΞΌ Ξ½).obs n = Kernel.const _ (ΞΌ n) := rfl +lemma obs_oblivious (n : β„•) : (Environment.oblivious ΞΌ Ξ½).obs n = Kernel.const _ (ΞΌ n) := rfl @[simp] -lemma feedback_obliviousEnv (n : β„•) : - (obliviousEnv ΞΌ Ξ½).feedback n = (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := rfl +lemma feedback_oblivious (n : β„•) : + (Environment.oblivious ΞΌ Ξ½).feedback n = (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := rfl @[simp] -lemma obs0_obliviousEnv : (obliviousEnv ΞΌ Ξ½).obs0 = ΞΌ 0 := rfl +lemma obsZero_oblivious : (Environment.oblivious ΞΌ Ξ½).obsZero = ΞΌ 0 := rfl @[simp] -lemma Ξ½0_obliviousEnv : (obliviousEnv ΞΌ Ξ½).Ξ½0 = Ξ½ 0 := by +lemma feedbackZero_oblivious : (Environment.oblivious ΞΌ Ξ½).feedbackZero = Ξ½ 0 := by ext p : 1 - rw [Environment.Ξ½0_def, Kernel.comap_apply, feedback_obliviousEnv, Kernel.comap_apply] + rw [Environment.feedbackZero_def, Kernel.comap_apply, feedback_oblivious, Kernel.comap_apply] -lemma stepKernel_obliviousEnv (alg : Algorithm π“ž 𝓐 𝓨) (n : β„•) : - stepKernel alg (obliviousEnv ΞΌ Ξ½) n +lemma stepKernel_oblivious (alg : Algorithm π“ž 𝓐 𝓨) (n : β„•) : + stepKernel alg (Environment.oblivious ΞΌ Ξ½) n = Kernel.const _ (ΞΌ n) βŠ—β‚– (alg.policy n βŠ—β‚– (Ξ½ n).comap (fun p ↦ (p.1.2, p.2)) (by fun_prop)) := rfl -instance : IsObliviousEnv (obliviousEnv ΞΌ Ξ½) where +instance : (Environment.oblivious ΞΌ Ξ½).IsOblivious where exists_obs_eq_const := ⟨μ, inferInstance, fun _ ↦ rfl⟩ exists_feedback_eq_comap := ⟨ν, inferInstance, fun _ ↦ rfl⟩ -/-- The law of the observations of `obliviousEnv ΞΌ Ξ½` is `ΞΌ`. The nonemptiness assumptions ensure -that there are histories of every length, so that the observation kernels determine `ΞΌ`. -/ +/-- The law of the observations of `Environment.oblivious ΞΌ Ξ½` is `ΞΌ`. The nonemptiness assumptions +ensure that there are histories of every length, so that the observation kernels determine `ΞΌ`. -/ @[simp] -lemma obsLaw_obliviousEnv [Nonempty 𝓐] [Nonempty 𝓨] (n : β„•) : - (obliviousEnv ΞΌ Ξ½).obsLaw n = ΞΌ n := by +lemma obsLaw_oblivious [Nonempty 𝓐] [Nonempty 𝓨] (n : β„•) : + (Environment.oblivious ΞΌ Ξ½).obsLaw n = ΞΌ n := by have : Nonempty π“ž := Measure.nonempty_of_neZero (ΞΌ n) - have h_eq := (obliviousEnv ΞΌ Ξ½).obs_eq_const_obsLaw n - rw [obs_obliviousEnv, Kernel.ext_iff] at h_eq + have h_eq := (Environment.oblivious ΞΌ Ξ½).obs_eq_const_obsLaw n + rw [obs_oblivious, Kernel.ext_iff] at h_eq simpa using (h_eq (Classical.arbitrary _)).symm @[simp] -lemma feedbackCondObsAction_obliviousEnv (n : β„•) : - (obliviousEnv ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ n := by +lemma feedbackCondObsAction_oblivious (n : β„•) : + (Environment.oblivious ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ n := by rcases isEmpty_or_nonempty π“ž with hπ“ž | hπ“ž Β· ext p : 1 exact hπ“ž.elim p.1 @@ -264,8 +264,8 @@ lemma feedbackCondObsAction_obliviousEnv (n : β„•) : Β· refine absurd (hΞ½ 0) ?_ simp only [Subsingleton.eq_zero Ξ½, Pi.zero_apply] exact Kernel.not_isMarkovKernel_zero - have h_eq := (obliviousEnv ΞΌ Ξ½).feedback_eq_comap_feedbackCondObsAction n - rw [feedback_obliviousEnv, Kernel.ext_iff] at h_eq + have h_eq := (Environment.oblivious ΞΌ Ξ½).feedback_eq_comap_feedbackCondObsAction n + rw [feedback_oblivious, Kernel.ext_iff] at h_eq ext p : 1 obtain ⟨o, a⟩ := p exact (h_eq ((Classical.arbitrary _, o), a)).symm @@ -279,42 +279,44 @@ variable {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] {Ξ½ : Kernel (π“ž Γ— /-- The stationary environment in which the observations have law `ΞΌ` and the feedback is drawn from `Ξ½` applied to the observation and the action, whatever the past history. -/ noncomputable -def stationaryEnv (ΞΌ : Measure π“ž) [IsProbabilityMeasure ΞΌ] (Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨) +def Environment.stationary (ΞΌ : Measure π“ž) [IsProbabilityMeasure ΞΌ] (Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨) [IsMarkovKernel Ξ½] : Environment π“ž 𝓐 𝓨 := - obliviousEnv (fun _ ↦ ΞΌ) (fun _ ↦ Ξ½) + Environment.oblivious (fun _ ↦ ΞΌ) (fun _ ↦ Ξ½) -lemma stationaryEnv_def : stationaryEnv ΞΌ Ξ½ = obliviousEnv (fun _ ↦ ΞΌ) (fun _ ↦ Ξ½) := rfl +lemma Environment.stationary_def : + Environment.stationary ΞΌ Ξ½ = Environment.oblivious (fun _ ↦ ΞΌ) (fun _ ↦ Ξ½) := rfl @[simp] -lemma obs_stationaryEnv (n : β„•) : (stationaryEnv ΞΌ Ξ½).obs n = Kernel.const _ ΞΌ := rfl +lemma obs_stationary (n : β„•) : (Environment.stationary ΞΌ Ξ½).obs n = Kernel.const _ ΞΌ := rfl @[simp] -lemma feedback_stationaryEnv (n : β„•) : - (stationaryEnv ΞΌ Ξ½).feedback n = Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := rfl +lemma feedback_stationary (n : β„•) : + (Environment.stationary ΞΌ Ξ½).feedback n = Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := rfl @[simp] -lemma obs0_stationaryEnv : (stationaryEnv ΞΌ Ξ½).obs0 = ΞΌ := rfl +lemma obsZero_stationary : (Environment.stationary ΞΌ Ξ½).obsZero = ΞΌ := rfl @[simp] -lemma Ξ½0_stationaryEnv : (stationaryEnv ΞΌ Ξ½).Ξ½0 = Ξ½ := Ξ½0_obliviousEnv +lemma feedbackZero_stationary : (Environment.stationary ΞΌ Ξ½).feedbackZero = Ξ½ := + feedbackZero_oblivious -lemma stepKernel_stationaryEnv (alg : Algorithm π“ž 𝓐 𝓨) (n : β„•) : - stepKernel alg (stationaryEnv ΞΌ Ξ½) n +lemma stepKernel_stationary (alg : Algorithm π“ž 𝓐 𝓨) (n : β„•) : + stepKernel alg (Environment.stationary ΞΌ Ξ½) n = Kernel.const _ ΞΌ βŠ—β‚– (alg.policy n βŠ—β‚– Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop)) := rfl -instance : IsObliviousEnv (stationaryEnv ΞΌ Ξ½) := - inferInstanceAs (IsObliviousEnv (obliviousEnv _ _)) +instance : (Environment.stationary ΞΌ Ξ½).IsOblivious := + inferInstanceAs (Environment.oblivious _ _).IsOblivious @[simp] -lemma obsLaw_stationaryEnv [Nonempty 𝓐] [Nonempty 𝓨] (n : β„•) : - (stationaryEnv ΞΌ Ξ½).obsLaw n = ΞΌ := - obsLaw_obliviousEnv n +lemma obsLaw_stationary [Nonempty 𝓐] [Nonempty 𝓨] (n : β„•) : + (Environment.stationary ΞΌ Ξ½).obsLaw n = ΞΌ := + obsLaw_oblivious n @[simp] -lemma feedbackCondObsAction_stationaryEnv (n : β„•) : - (stationaryEnv ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ := - feedbackCondObsAction_obliviousEnv n +lemma feedbackCondObsAction_stationary (n : β„•) : + (Environment.stationary ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ := + feedbackCondObsAction_oblivious n end Stationary @@ -327,11 +329,11 @@ variable {Ξ½ : β„• β†’ Kernel 𝓐 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] noncomputable def Environment.banditSeq (Ξ½ : β„• β†’ Kernel 𝓐 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] : Environment Unit 𝓐 𝓨 := - obliviousEnv (fun _ ↦ Measure.dirac ()) (fun n ↦ (Ξ½ n).prodMkLeft Unit) + Environment.oblivious (fun _ ↦ Measure.dirac ()) (fun n ↦ (Ξ½ n).prodMkLeft Unit) lemma Environment.banditSeq_def : Environment.banditSeq Ξ½ - = obliviousEnv (fun _ ↦ Measure.dirac ()) (fun n ↦ (Ξ½ n).prodMkLeft Unit) := + = Environment.oblivious (fun _ ↦ Measure.dirac ()) (fun n ↦ (Ξ½ n).prodMkLeft Unit) := rfl @[simp] @@ -342,13 +344,14 @@ lemma obs_banditSeq (n : β„•) : lemma feedback_banditSeq (n : β„•) : (Environment.banditSeq Ξ½).feedback n = (Ξ½ n).prodMkLeft _ := rfl @[simp] -lemma obs0_banditSeq : (Environment.banditSeq Ξ½).obs0 = Measure.dirac () := rfl +lemma obsZero_banditSeq : (Environment.banditSeq Ξ½).obsZero = Measure.dirac () := rfl @[simp] -lemma Ξ½0_banditSeq : (Environment.banditSeq Ξ½).Ξ½0 = (Ξ½ 0).prodMkLeft Unit := Ξ½0_obliviousEnv +lemma feedbackZero_banditSeq : (Environment.banditSeq Ξ½).feedbackZero = (Ξ½ 0).prodMkLeft Unit := + feedbackZero_oblivious -instance : IsObliviousEnv (Environment.banditSeq Ξ½) := - inferInstanceAs (IsObliviousEnv (obliviousEnv _ _)) +instance : (Environment.banditSeq Ξ½).IsOblivious := + inferInstanceAs (Environment.oblivious _ _).IsOblivious @[simp] lemma obsLaw_banditSeq (n : β„•) : (Environment.banditSeq Ξ½).obsLaw n = Measure.dirac () := @@ -357,7 +360,7 @@ lemma obsLaw_banditSeq (n : β„•) : (Environment.banditSeq Ξ½).obsLaw n = Measure @[simp] lemma feedbackCondObsAction_banditSeq (n : β„•) : (Environment.banditSeq Ξ½).feedbackCondObsAction n = (Ξ½ n).prodMkLeft Unit := - feedbackCondObsAction_obliviousEnv n + feedbackCondObsAction_oblivious n end BanditSeq @@ -369,10 +372,10 @@ variable {Ξ½ : Kernel 𝓐 𝓨} [IsMarkovKernel Ξ½] applied to the action, whatever the past history: a stochastic bandit. -/ noncomputable def Environment.bandit (Ξ½ : Kernel 𝓐 𝓨) [IsMarkovKernel Ξ½] : Environment Unit 𝓐 𝓨 := - stationaryEnv (Measure.dirac ()) (Ξ½.prodMkLeft Unit) + Environment.stationary (Measure.dirac ()) (Ξ½.prodMkLeft Unit) lemma Environment.bandit_def : - Environment.bandit Ξ½ = stationaryEnv (Measure.dirac ()) (Ξ½.prodMkLeft Unit) := rfl + Environment.bandit Ξ½ = Environment.stationary (Measure.dirac ()) (Ξ½.prodMkLeft Unit) := rfl lemma Environment.bandit_eq_banditSeq : Environment.bandit Ξ½ = Environment.banditSeq fun _ ↦ Ξ½ := rfl @@ -389,13 +392,14 @@ lemma stepKernel_bandit (alg : Algorithm Unit 𝓐 𝓨) (n : β„•) : rw [stepKernel_def, obs_bandit, feedback_bandit] @[simp] -lemma obs0_bandit : (Environment.bandit Ξ½).obs0 = Measure.dirac () := rfl +lemma obsZero_bandit : (Environment.bandit Ξ½).obsZero = Measure.dirac () := rfl @[simp] -lemma Ξ½0_bandit : (Environment.bandit Ξ½).Ξ½0 = Ξ½.prodMkLeft Unit := Ξ½0_obliviousEnv +lemma feedbackZero_bandit : (Environment.bandit Ξ½).feedbackZero = Ξ½.prodMkLeft Unit := + feedbackZero_oblivious -instance : IsObliviousEnv (Environment.bandit Ξ½) := - inferInstanceAs (IsObliviousEnv (obliviousEnv _ _)) +instance : (Environment.bandit Ξ½).IsOblivious := + inferInstanceAs (Environment.oblivious _ _).IsOblivious @[simp] lemma obsLaw_bandit (n : β„•) : (Environment.bandit Ξ½).obsLaw n = Measure.dirac () := @@ -404,7 +408,7 @@ lemma obsLaw_bandit (n : β„•) : (Environment.bandit Ξ½).obsLaw n = Measure.dirac @[simp] lemma feedbackCondObsAction_bandit (n : β„•) : (Environment.bandit Ξ½).feedbackCondObsAction n = Ξ½.prodMkLeft Unit := - feedbackCondObsAction_obliviousEnv n + feedbackCondObsAction_oblivious n end Bandit @@ -417,38 +421,38 @@ variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} {alg : Algorithm π“ž 𝓐 𝓨 {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} /-- The observation at time `n` has law `ΞΌ n`. -/ -lemma hasLaw_obs_obliviousEnv {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] +lemma hasLaw_obs_oblivious {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] {Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] - (h : IsAlgEnvSeq O A Y alg (obliviousEnv ΞΌ Ξ½) P) (n : β„•) : + (h : IsAlgEnvSeq O A Y alg (Environment.oblivious ΞΌ Ξ½) P) (n : β„•) : HasLaw (O n) (ΞΌ n) P := by have h' := h.hasCondDistrib_obs n - rw [obs_obliviousEnv] at h' + rw [obs_oblivious] at h' exact h'.hasLaw_of_const /-- The conditional distribution of the feedback at time `n` given the observation and the action at time `n` is `Ξ½ n`. -/ -lemma hasCondDistrib_feedback_obliviousEnv {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] +lemma hasCondDistrib_feedback_oblivious {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] {Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] - (h : IsAlgEnvSeq O A Y alg (obliviousEnv ΞΌ Ξ½) P) (n : β„•) : + (h : IsAlgEnvSeq O A Y alg (Environment.oblivious ΞΌ Ξ½) P) (n : β„•) : HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) (Ξ½ n) P := by have h' := h.hasCondDistrib_feedback n - rw [feedback_obliviousEnv] at h' + rw [feedback_oblivious] at h' exact h'.comp_right /-- The observation at time `n` has law `ΞΌ`. -/ -lemma hasLaw_obs_stationaryEnv {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] +lemma hasLaw_obs_stationary {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] {Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨} [IsMarkovKernel Ξ½] - (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΌ Ξ½) P) (n : β„•) : + (h : IsAlgEnvSeq O A Y alg (Environment.stationary ΞΌ Ξ½) P) (n : β„•) : HasLaw (O n) ΞΌ P := - hasLaw_obs_obliviousEnv h n + hasLaw_obs_oblivious h n /-- The conditional distribution of the feedback at time `n` given the observation and the action at time `n` is `Ξ½`. -/ -lemma hasCondDistrib_feedback_stationaryEnv {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] +lemma hasCondDistrib_feedback_stationary {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] {Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨} [IsMarkovKernel Ξ½] - (h : IsAlgEnvSeq O A Y alg (stationaryEnv ΞΌ Ξ½) P) (n : β„•) : + (h : IsAlgEnvSeq O A Y alg (Environment.stationary ΞΌ Ξ½) P) (n : β„•) : HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) Ξ½ P := - hasCondDistrib_feedback_obliviousEnv h n + hasCondDistrib_feedback_oblivious h n end General @@ -465,7 +469,7 @@ lemma hasCondDistrib_feedback_banditSeq {Ξ½ : β„• β†’ Kernel 𝓐 𝓨} [βˆ€ n, (h : IsAlgEnvSeq O A Y alg (Environment.banditSeq Ξ½) P) (n : β„•) : HasCondDistrib (Y n) (A n) (Ξ½ n) P := by have h' : HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) ((Ξ½ n).prodMkLeft Unit) P := - hasCondDistrib_feedback_obliviousEnv h n + hasCondDistrib_feedback_oblivious h n exact h'.comp_right /-- The conditional distribution of the feedback at time `n` given the action at time `n` is `Ξ½`. -/ @@ -487,7 +491,7 @@ lemma hasLaw_feedback_cond_bandit (h : IsAlgEnvSeq O A Y alg (Environment.bandit (hsb : βˆ€ u ∈ s, u.2 = b) (hP : P ((fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s) β‰  0) : HasLaw (Y n) (Ξ½ b) P[|(fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s] := by - simpa using IsObliviousEnv.hasLaw_feedback_cond h n hs (b := ((), b)) + simpa using Environment.IsOblivious.hasLaw_feedback_cond h n hs (b := ((), b)) (fun u hu ↦ by simp [hsb u hu]) hP /-- Conditionally on an event determined by the history before time `n` and the action at time @@ -499,7 +503,7 @@ lemma indepFun_history_action_feedback_cond_bandit (hsb : βˆ€ u ∈ s, u.2 = b) : (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) βŸ‚α΅’[P[|(fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s]] Y n := - IsObliviousEnv.indepFun_history_action_feedback_cond h n hs (b := ((), b)) + Environment.IsOblivious.indepFun_history_action_feedback_cond h n hs (b := ((), b)) fun u hu ↦ by simp [hsb u hu] /-- The feedback at time `n` is conditionally independent of the history before time `n` From fbe84023e403c18ad5b9a08e6c1284a4697a10bb Mon Sep 17 00:00:00 2001 From: Remy Degenne Date: Sat, 3 Oct 2026 15:24:35 +0200 Subject: [PATCH 6/7] add Environment.IsStationary --- .../SequentialLearning/EvaluationEnv.lean | 8 +- .../SequentialLearning/Means.lean | 5 + .../SequentialLearning/README.md | 2 +- .../SequentialLearning/StationaryEnv.lean | 170 +++++++++++++++--- 4 files changed, 155 insertions(+), 30 deletions(-) diff --git a/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean b/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean index 3d94838f..d60e522a 100644 --- a/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean +++ b/LeanMachineLearning/SequentialLearning/EvaluationEnv.lean @@ -24,7 +24,8 @@ to change at every time step, while the second one uses a fixed function at ever kernel that evaluates a fixed measurable function `f` at the chosen action. They both satisfy the typeclasses `Environment.IsOblivious` and -`Environment.HasDeterministicFeedback`. +`Environment.HasDeterministicFeedback`, and `Environment.eval f hf` is also +`Environment.IsStationary`. ## Main statements @@ -108,6 +109,11 @@ end OnlineEvalEnv noncomputable def Environment.eval (f : 𝓐 β†’ 𝓨) (hf : Measurable f) := Environment.evalSeq (fun _ ↦ f) (fun _ ↦ hf) +instance : (Environment.eval f hf).IsStationary where + exists_obs_eq_const := ⟨Measure.dirac (), inferInstance, fun _ ↦ rfl⟩ + exists_feedback_eq_comap := + ⟨(Kernel.deterministic f hf).prodMkLeft Unit, inferInstance, fun _ ↦ rfl⟩ + instance : (Environment.eval f hf).IsOblivious := by unfold Environment.eval; infer_instance instance : (Environment.eval f hf).HasDeterministicFeedback := by diff --git a/LeanMachineLearning/SequentialLearning/Means.lean b/LeanMachineLearning/SequentialLearning/Means.lean index dc1418e7..62460fcb 100644 --- a/LeanMachineLearning/SequentialLearning/Means.lean +++ b/LeanMachineLearning/SequentialLearning/Means.lean @@ -75,6 +75,11 @@ lemma means_of_isOblivious [env.IsOblivious] (O : β„• β†’ Ξ© β†’ π“ž) (A : β„• simp [Environment.means, Environment.measure, env.feedback_eq_comap_feedbackCondObsAction, Kernel.comap_apply] +lemma means_of_isStationary [env.IsStationary] (O : β„• β†’ Ξ© β†’ π“ž) (A : β„• β†’ Ξ© β†’ 𝓐) + (Y : β„• β†’ Ξ© β†’ 𝓨) (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : + env.means O A Y k n Ο‰ = (env.feedbackZero (O n Ο‰, k))[id] := by + rw [means_of_isOblivious, env.feedbackCondObsAction_eq_feedbackZero] + lemma means_oblivious (ΞΌ : β„• β†’ Measure π“ž) [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] (Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨) [βˆ€ n, IsMarkovKernel (Ξ½ n)] (k : 𝓐) (n : β„•) (Ο‰ : Ξ©) : (Environment.oblivious ΞΌ Ξ½).means O A Y k n Ο‰ = (Ξ½ n (O n Ο‰, k))[id] := by simp diff --git a/LeanMachineLearning/SequentialLearning/README.md b/LeanMachineLearning/SequentialLearning/README.md index f21fcca3..e34f5e8c 100644 --- a/LeanMachineLearning/SequentialLearning/README.md +++ b/LeanMachineLearning/SequentialLearning/README.md @@ -62,7 +62,7 @@ All for obs, no action for feedback: `Environment.FeedbackIgnoresAction` (MISSIN No history for obs and feedback: `Environment.IsOblivious`, `Environment.oblivious`. -No time, no history for obs and feedback: `Environment.IsStationary` (MISSING), `Environment.stationary`. +No time, no history for obs and feedback: `Environment.IsStationary`, `Environment.stationary`. No time, last round of history for obs, not history for feedback: `Environment.IsMarkov` (MISSING), `Environment.markov` (MISSING). diff --git a/LeanMachineLearning/SequentialLearning/StationaryEnv.lean b/LeanMachineLearning/SequentialLearning/StationaryEnv.lean index da693d3b..dcd5571d 100644 --- a/LeanMachineLearning/SequentialLearning/StationaryEnv.lean +++ b/LeanMachineLearning/SequentialLearning/StationaryEnv.lean @@ -15,22 +15,27 @@ An oblivious environment is an environment in which the distributions of the obs the feedback do not depend on the past history: at time `n`, the observation has law `env.obsLaw n`, and the feedback depends only on the current observation and action, through the Markov kernel `env.feedbackCondObsAction n`. -If there are no observations (`π“ž = Unit`) and the kernel that gives the distribution of the -feedback given the action is the same at every time step, then we say that the environment is -stationary. +A stationary environment is an oblivious environment in which these laws do not depend on time +either: every observation has law `env.obsZero`, and the feedback is drawn from the Markov kernel +`env.feedbackZero` applied to the current observation and action. ## Main definitions -We define a `Prop`-valued typeclass `Environment.IsOblivious` to express that an environment is -oblivious, and we define constructors for oblivious environments, with and without observations. +We define `Prop`-valued typeclasses `Environment.IsOblivious` and `Environment.IsStationary` to +express that an environment is oblivious or stationary, and we define constructors for oblivious +and stationary environments, with and without observations. -Typeclass and related definitions: +Typeclasses and related definitions: * `Environment.IsOblivious env`: the environment `env` is oblivious. * `Environment.obsLaw env n`: the law of the observation at time `n` in an oblivious environment `env`. * `Environment.feedbackCondObsAction env n`: the kernel representing the conditional distribution of the feedback given the observation and the action at time `n` in an oblivious environment `env`. +* `Environment.IsStationary env`: the environment `env` is stationary. A stationary environment is + oblivious, and its laws are described by the time zero accessors `env.obsZero` and + `env.feedbackZero` (see `Environment.obs_eq_const_obsZero` and + `Environment.feedback_eq_comap_feedbackZero`). Constructors for oblivious environments: * `Environment.oblivious ΞΌ Ξ½`: the oblivious environment in which the observation at time `n` has @@ -105,6 +110,35 @@ lemma feedbackZero_eq_feedbackCondObsAction (env : Environment π“ž 𝓐 𝓨) [ rw [Environment.feedbackZero_def, Kernel.comap_apply, feedback_eq_comap_feedbackCondObsAction, Kernel.comap_apply] +lemma obsLaw_eq_of_obs_eq_const [Nonempty 𝓐] [Nonempty 𝓨] + (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] {n : β„•} {ΞΌ : Measure π“ž} [IsProbabilityMeasure ΞΌ] + (h : env.obs n = Kernel.const _ ΞΌ) : + env.obsLaw n = ΞΌ := by + have : Nonempty π“ž := Measure.nonempty_of_neZero ΞΌ + have h_eq := env.obs_eq_const_obsLaw n + rw [h, Kernel.ext_iff] at h_eq + simpa using (h_eq (Classical.arbitrary _)).symm + +lemma feedbackCondObsAction_eq_of_feedback_eq (env : Environment π“ž 𝓐 𝓨) [env.IsOblivious] + {n : β„•} {Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨} [hΞ½ : IsMarkovKernel Ξ½] + (h : env.feedback n = Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop)) : + env.feedbackCondObsAction n = Ξ½ := by + rcases isEmpty_or_nonempty π“ž with hπ“ž | hπ“ž + Β· ext p : 1 + exact hπ“ž.elim p.1 + rcases isEmpty_or_nonempty 𝓐 with h𝓐 | h𝓐 + Β· ext p : 1 + exact h𝓐.elim p.2 + rcases isEmpty_or_nonempty 𝓨 with h𝓨 | h𝓨 + Β· refine absurd hΞ½ ?_ + rw [Subsingleton.eq_zero Ξ½] + exact Kernel.not_isMarkovKernel_zero + have h_eq := env.feedback_eq_comap_feedbackCondObsAction n + rw [h, Kernel.ext_iff] at h_eq + ext p : 1 + obtain ⟨o, a⟩ := p + exact (h_eq ((Classical.arbitrary _, o), a)).symm + end Environment namespace Environment.IsOblivious @@ -202,10 +236,100 @@ lemma condIndepFun_feedback_history_obs_action [StandardBorelSpace Ξ©] end Environment.IsOblivious +/-- An environment is stationary if it is oblivious and its laws do not depend on time: the +observations have a fixed law, and the feedback depends only on the current observation and action, +through a fixed Markov kernel. -/ +class Environment.IsStationary (env : Environment π“ž 𝓐 𝓨) : Prop where + exists_obs_eq_const : βˆƒ ΞΌ : Measure π“ž, IsProbabilityMeasure ΞΌ ∧ βˆ€ n, env.obs n = Kernel.const _ ΞΌ + exists_feedback_eq_comap : βˆƒ Ξ½ : Kernel (π“ž Γ— 𝓐) 𝓨, IsMarkovKernel Ξ½ ∧ + βˆ€ n, env.feedback n = Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) + +namespace Environment + +/-- In a stationary environment, the observation at every time has law `env.obsZero`. -/ +lemma obs_eq_const_obsZero (env : Environment π“ž 𝓐 𝓨) [h : env.IsStationary] (n : β„•) : + env.obs n = Kernel.const _ env.obsZero := by + obtain ⟨μ, -, hμ⟩ := h.exists_obs_eq_const + rw [hΞΌ n, obsZero_def, hΞΌ 0, Kernel.const_apply] + +/-- In a stationary environment, the feedback at every time is drawn from `env.feedbackZero` applied +to the current observation and action. -/ +lemma feedback_eq_comap_feedbackZero (env : Environment π“ž 𝓐 𝓨) [h : env.IsStationary] (n : β„•) : + env.feedback n = env.feedbackZero.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) := by + obtain ⟨ν, -, hν⟩ := h.exists_feedback_eq_comap + have hΞ½0 : env.feedbackZero = Ξ½ := by + ext p : 1 + rw [feedbackZero_def, Kernel.comap_apply, hΞ½ 0, Kernel.comap_apply] + rw [hΞ½ n, hΞ½0] + +instance (env : Environment π“ž 𝓐 𝓨) [env.IsStationary] : env.IsOblivious where + exists_obs_eq_const := ⟨fun _ ↦ env.obsZero, fun _ ↦ inferInstance, env.obs_eq_const_obsZero⟩ + exists_feedback_eq_comap := + ⟨fun _ ↦ env.feedbackZero, fun _ ↦ inferInstance, env.feedback_eq_comap_feedbackZero⟩ + +/-- In a stationary environment, the law of the observation at time `n` is `env.obsZero`. -/ +lemma obsLaw_eq_obsZero (env : Environment π“ž 𝓐 𝓨) [env.IsStationary] [Nonempty 𝓐] [Nonempty 𝓨] + (n : β„•) : + env.obsLaw n = env.obsZero := + env.obsLaw_eq_of_obs_eq_const (env.obs_eq_const_obsZero n) + +/-- In a stationary environment, the conditional distribution of the feedback given the observation +and the action at time `n` is `env.feedbackZero`. -/ +lemma feedbackCondObsAction_eq_feedbackZero (env : Environment π“ž 𝓐 𝓨) [env.IsStationary] + (n : β„•) : + env.feedbackCondObsAction n = env.feedbackZero := + env.feedbackCondObsAction_eq_of_feedback_eq (env.feedback_eq_comap_feedbackZero n) + +end Environment + +namespace Environment.IsStationary + +variable {Ξ© : Type*} {mΞ© : MeasurableSpace Ξ©} + {alg : Algorithm π“ž 𝓐 𝓨} {env : Environment π“ž 𝓐 𝓨} {P : Measure Ξ©} + {O : β„• β†’ Ξ© β†’ π“ž} {A : β„• β†’ Ξ© β†’ 𝓐} {Y : β„• β†’ Ξ© β†’ 𝓨} + +/-- The observation at time `n` has law `env.obsZero`. -/ +lemma hasLaw_obs [IsProbabilityMeasure P] [env.IsStationary] + (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + HasLaw (O n) env.obsZero P := by + have h' := h.hasCondDistrib_obs n + rw [env.obs_eq_const_obsZero] at h' + exact h'.hasLaw_of_const + +variable [IsFiniteMeasure P] + +lemma hasCondDistrib_feedback_history_action [env.IsStationary] + (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + HasCondDistrib (Y n) (fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) + (env.feedbackZero.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop) + : Kernel ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐) 𝓨) P := by + rw [← env.feedback_eq_comap_feedbackZero] + exact h.hasCondDistrib_feedback n + +/-- The conditional distribution of the feedback at time `n` given the observation and the action +at time `n` is `env.feedbackZero`. -/ +lemma hasCondDistrib_feedback [env.IsStationary] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) : + HasCondDistrib (Y n) (fun Ο‰ ↦ (O n Ο‰, A n Ο‰)) env.feedbackZero P := + (hasCondDistrib_feedback_history_action h n).comp_right + +/-- Conditionally on an event determined by the history before time `n`, the observation and the +action at time `n`, on which the observation-action pair is equal to `b`, the feedback at time `n` +has law `env.feedbackZero b`. -/ +lemma hasLaw_feedback_cond [env.IsStationary] (h : IsAlgEnvSeq O A Y alg env P) (n : β„•) + {s : Set ((Hist π“ž 𝓐 𝓨 n Γ— π“ž) Γ— 𝓐)} (hs : MeasurableSet s) {b : π“ž Γ— 𝓐} + (hsb : βˆ€ u ∈ s, (u.1.2, u.2) = b) + (hP : P ((fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s) β‰  0) : + HasLaw (Y n) (env.feedbackZero b) + P[|(fun Ο‰ ↦ ((history O A Y n Ο‰, O n Ο‰), A n Ο‰)) ⁻¹' s] := by + rw [← env.feedbackCondObsAction_eq_feedbackZero n] + exact IsOblivious.hasLaw_feedback_cond h n hs hsb hP + +end Environment.IsStationary + section Oblivious variable {ΞΌ : β„• β†’ Measure π“ž} [βˆ€ n, IsProbabilityMeasure (ΞΌ n)] - {Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨} [hΞ½ : βˆ€ n, IsMarkovKernel (Ξ½ n)] + {Ξ½ : β„• β†’ Kernel (π“ž Γ— 𝓐) 𝓨} [βˆ€ n, IsMarkovKernel (Ξ½ n)] /-- The oblivious environment in which the observation at time `n` has law `ΞΌ n` and the feedback at time `n` is drawn from `Ξ½ n` applied to the observation and the action at time `n`, whatever the @@ -245,30 +369,13 @@ instance : (Environment.oblivious ΞΌ Ξ½).IsOblivious where ensure that there are histories of every length, so that the observation kernels determine `ΞΌ`. -/ @[simp] lemma obsLaw_oblivious [Nonempty 𝓐] [Nonempty 𝓨] (n : β„•) : - (Environment.oblivious ΞΌ Ξ½).obsLaw n = ΞΌ n := by - have : Nonempty π“ž := Measure.nonempty_of_neZero (ΞΌ n) - have h_eq := (Environment.oblivious ΞΌ Ξ½).obs_eq_const_obsLaw n - rw [obs_oblivious, Kernel.ext_iff] at h_eq - simpa using (h_eq (Classical.arbitrary _)).symm + (Environment.oblivious ΞΌ Ξ½).obsLaw n = ΞΌ n := + Environment.obsLaw_eq_of_obs_eq_const _ rfl @[simp] lemma feedbackCondObsAction_oblivious (n : β„•) : - (Environment.oblivious ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ n := by - rcases isEmpty_or_nonempty π“ž with hπ“ž | hπ“ž - Β· ext p : 1 - exact hπ“ž.elim p.1 - rcases isEmpty_or_nonempty 𝓐 with h𝓐 | h𝓐 - Β· ext p : 1 - exact h𝓐.elim p.2 - rcases isEmpty_or_nonempty 𝓨 with h𝓨 | h𝓨 - Β· refine absurd (hΞ½ 0) ?_ - simp only [Subsingleton.eq_zero Ξ½, Pi.zero_apply] - exact Kernel.not_isMarkovKernel_zero - have h_eq := (Environment.oblivious ΞΌ Ξ½).feedback_eq_comap_feedbackCondObsAction n - rw [feedback_oblivious, Kernel.ext_iff] at h_eq - ext p : 1 - obtain ⟨o, a⟩ := p - exact (h_eq ((Classical.arbitrary _, o), a)).symm + (Environment.oblivious ΞΌ Ξ½).feedbackCondObsAction n = Ξ½ n := + Environment.feedbackCondObsAction_eq_of_feedback_eq _ rfl end Oblivious @@ -305,6 +412,10 @@ lemma stepKernel_stationary (alg : Algorithm π“ž 𝓐 𝓨) (n : β„•) : = Kernel.const _ ΞΌ βŠ—β‚– (alg.policy n βŠ—β‚– Ξ½.comap (fun p ↦ (p.1.2, p.2)) (by fun_prop)) := rfl +instance : (Environment.stationary ΞΌ Ξ½).IsStationary where + exists_obs_eq_const := ⟨μ, inferInstance, fun _ ↦ rfl⟩ + exists_feedback_eq_comap := ⟨ν, inferInstance, fun _ ↦ rfl⟩ + instance : (Environment.stationary ΞΌ Ξ½).IsOblivious := inferInstanceAs (Environment.oblivious _ _).IsOblivious @@ -398,6 +509,9 @@ lemma obsZero_bandit : (Environment.bandit Ξ½).obsZero = Measure.dirac () := rfl lemma feedbackZero_bandit : (Environment.bandit Ξ½).feedbackZero = Ξ½.prodMkLeft Unit := feedbackZero_oblivious +instance : (Environment.bandit Ξ½).IsStationary := + inferInstanceAs (Environment.stationary _ _).IsStationary + instance : (Environment.bandit Ξ½).IsOblivious := inferInstanceAs (Environment.oblivious _ _).IsOblivious From f6d21903a388606e90734273cff06968ba417a27 Mon Sep 17 00:00:00 2001 From: Remy Degenne Date: Sat, 3 Oct 2026 15:28:51 +0200 Subject: [PATCH 7/7] add note about MISSING names --- LeanMachineLearning/SequentialLearning/README.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/LeanMachineLearning/SequentialLearning/README.md b/LeanMachineLearning/SequentialLearning/README.md index e34f5e8c..eb0cecbf 100644 --- a/LeanMachineLearning/SequentialLearning/README.md +++ b/LeanMachineLearning/SequentialLearning/README.md @@ -27,6 +27,10 @@ structure Environment (π“ž 𝓐 𝓨 : Type*) [MeasurableSpace π“ž] [Measurabl In many applications, some of those kernels are deterministic, or do not depend on some of their inputs. We detail here the naming conventions for the various constructors, predicates and accessors that are used in the library. +NOTE: some names described below are marked as MISSING because they are not yet implemented in the library. +They may never be implemented if we don't need them. +If you need one of them, treat the name here as a recommendation that you may want to use. + Generic constructions live in the `Algorithm` and `Environment` namespaces. Predicates live in the `Algorithm` and `Environment` namespaces: they are classes when they carry an accessor, and `Prop` definitions otherwise. Accessors are namespaced so that dot notation works.