Arturas Juodis, Ovidijus Stauskas, Sander Tromp
Mon 28 Sep
econ.EM
pdf html
- interest Causal inference and identification in econometrics 0.92
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.91
- reading inference, bootstrap, large, dynamic
Abstract
We study recursive-design wild bootstrap inference for dynamic panel data models with unobserved common factors estimated by Common Correlated Effects. In the large N,T setting, the bootstrap reproduces the biased limiting distribution in pure autoregressive models, but fails to capture all bias and factor-estimation variance components in models with additional regressors, particularly under weak exogeneity. We trace this failure to holding regressors fixed across bootstrap replications. We propose to combine bootstrap procedure with available bias-correction methods to conduct adjusted inference. Monte Carlo evidence shows substantial improvements over conventional strategies of using bias-correction paired with cross-sectional bootstrap methods.
Shibo Yu, Yan Chen, Jin-Hong Du, Guodong Li
Thu 24 Sep
stat.ME
pdf html
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.93
- interest Causal inference and identification in econometrics 0.90
- reading semiparametric, series, causal effects, observational
Abstract
In observational time series, statistical inference for dynamic causal effects of a one-time intervention across horizons is complicated by high-dimensional observed pre-treatment information, unmeasured confounding, and serial dependence. To address these challenges, we develop a semiparametric framework for inference from a single serially dependent time series, integrating debiased machine learning with instrumental variables through buffered block cross-fitting. Under geometric beta-mixing, we derive non-asymptotic bounds on estimation error, asymptotic normality at each fixed horizon, and feasible inference that accommodates serial dependence. We further show how learner-specific prediction guarantees under temporal dependence can be used to verify the nuisance-rate conditions required for orthogonal inference. In a monetary-policy application with 468 months and 1464 lagged FRED-MD controls, we show an instrumented policy tightening lowers housing starts at medium horizons, with sensitivity analyses that support the finding.
Spandan Ghose Chowdhury
Thu 24 Sep
stat.MEcs.AIcs.LG
pdf html
- interest Causal inference and identification in econometrics 0.92
- reading observational, did, just, researchers
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.89
Abstract
Information systems (IS) researchers increasingly use quasi-experimental methods such as difference-in-differences (DiD) and instrumental variables (IV) to recover causal effects from observational panel data. Power calculations that justify these designs assume i.i.d. errors, but the deeper problem is what even a cluster-robust calculator cannot see. We report a Monte Carlo study over 9837 parameter conditions (approx 9.8 million datasets) and decompose the planned-versus-achieved power gap. The serial-correlation component is recoverable by an AR(1)-aware calculator when rho is known, and partially when rho must be estimated from short pre-periods, but panel attrition, staggered-adoption bias, and parallel-trends pretesting are captured by no closed-form formula; exogenous attrition alone costs approx 8 to 11 percentage points at the few-hundred-to-thousand sample sizes IS studies use. Treatment-correlated, outcome-dependent attrition instead induces bias, not just power loss. For IV, holding first-stage F fixed, larger N neither raises power nor curbs exclusion bias, though with a fixed instrument more data does sharpen the first stage, so identification rests on instrument strength, not sample size.
Marcelo J. Moreira, Whitney K. Newey, Mahrad Sharifvaghefi
Wed 30 Sep
econ.EM
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.87
- interest Causal inference and identification in econometrics 0.89
- reading endogenous, theory, gmm, gives
Abstract
This paper characterizes the global minimum of the continuously updating generalized method of moments (CU-GMM) objective in linear instrumental variables models. We allow optimal weighting matrices under heteroskedasticity, autocorrelation, or clustering. We show that the objective is a ratio of polynomials. For one endogenous regressor, stationary points of CU-GMM objective function are real eigenvalues of a companion matrix. Comparing their objective values with the value at infinity gives the global minimum, extending the classical eigenvector approach to limited information maximum likelihood. With multiple endogenous regressors, algebraic elimination and checks for real solutions identify the minimum among finitely many candidate objective values, including boundary values. Galois theory rules out general formulas by radicals even with one endogenous regressor and two instruments, while numerical root finding remains possible. Finding the global minimum allows us to compute overidentification and likelihood ratio tests, including the conditional likelihood ratio (CLR) test.
Grace V. Ringlein, Trang Q. Nguyen, Elizabeth A. Stuart, Harsh Parikh
Wed 30 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.93
- reading equations, treatment, functions, causal
- interest Causal inference and identification in econometrics 0.90
Abstract
Proximal causal inference identifies causal effects in the presence of unmeasured confounding by drawing on two sets of proxy variables. Identification typically utilizes bridge functions, defined as solutions to integral equations. However, the mechanism by which fitted bridge functions correct for confounding bias remains opaque, offering little to interpret, inspect or stress-test. We show that the defining equation of a treatment bridge function is already a balance condition, with a cross-proxy form: the weights are functions of the treatment confounding proxies and covariates, and they balance the outcome confounding proxies and covariates (i.e., reweighting the distribution in a particular treatment arm to match the distribution across treatment arms). Several existing identification paths via a treatment bridge function can then be interpreted as providing conditions under which balance on the outcome proxies implies balance on the unobserved confounders, which we call balance propagation. Leveraging this framing, we show that estimation of the treatment bridge function is a type of balancing weight estimation. Finally, we show that an outcome-weighted estimator form can also be obtained for a large class of proximal estimators, including those that utilize an outcome bridge function. Using this framing, we provide conditions under which common estimators are numerically equivalent.
Adam Bilchouris, Andriy Olenko
Wed 30 Sep
stat.MEmath.PRstat.CO
pdf html
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.89
- interest Causal inference and identification in econometrics 0.92
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.85
Abstract
The paper introduces the R package CovEsts, which implements several nonparametric estimators for the autocovariance function. First, it presents the theoretical foundations of the implemented estimators, their properties, assumptions, and potential limitations. Next, it outlines the structure of the package and its key functions, including several methods for estimating autocovariance functions, constructing corresponding bootstrap confidence regions, and correcting the provided estimators. The package also includes diagnostic tools, such as several metrics for comparing estimators, and additional functions for broader use. The article illustrates a high degree of flexibility of the package in the selection of function parameters and the tuning of the estimators. Applications of selected estimators and package functions are illustrated using simulated data, yearly sunspot counts, and US unemployment increments data.
Yaroslav Korobka, Vira Semenova
Wed 30 Sep
econ.EM
pdf html
- reading wage, nuisance, theory, partially identified
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.86
- favorites near “When Did Growth Begin? New Estimates of Productivity Growth in England from 1250 to 1870” 0.81
Abstract
We study estimation and inference for a partially identified parameter whose identified set depends on a first-stage nuisance parameter that must itself be estimated. Combining the criterion-function approach with the theory of Neyman-orthogonal moments that underlies double/debiased machine learning, we propose a two-step procedure: the point-identified nuisance is estimated by flexible machine-learning methods, and the set-identified target is recovered as a level set of a sample criterion built from orthogonal moment inequalities with cross-fitting. When the contour level is bounded, we show that the resulting set estimator converges in Hausdorff distance at the parametric rate of the infeasible criterion built on the true nuisance. We further develop a subsampling procedure that delivers asymptotically valid coverage, provided the product of the first-stage estimation errors is $o(N^{-1/2})$. We illustrate the method on bounds for the wage distribution and the interquantile range under selection into employment and on the gender wage gap with an interval-censored wage. The empirical application studies the gender wage gap using the March supplement of the 2015 Current Population Survey.
Xiaojun Song, Haojiao Zhao
Thu 24 Sep
econ.EM
pdf html
- interest Causal inference and identification in econometrics 0.91
- favorites near “Difference-in-Differences with multiple time periods” 0.89
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.85
Abstract
We propose a nonparametric test for unobserved treatment effect heterogeneity in regression discontinuity designs. Under the null of no unobserved heterogeneity, a transformed outcome that imputes treated potential outcomes for untreated units must have a continuous conditional distribution at the cutoff. We convert this implication into an integrated conditional-moment restriction using characteristic functions, thereby allowing the conditional local average treatment effect to be an unrestricted function of covariates. We derive the asymptotic distribution of the test statistics via a $U$-process and establish the validity of a multiplier bootstrap procedure for calculating critical values. Monte Carlo experiments show well-controlled size and increasing power. Two empirical applications illustrate how the test distinguishes between heterogeneity explained by observables and that explained by unobserved factors.
Mikihito Nishi, Takashi Yamagata
Tue 29 Sep
stat.MEecon.EM
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.83
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.85
- interest Macroeconometrics and time series 0.90
Abstract
This paper develops a general two-point dependent wild bootstrap (DWB) for weakly dependent estimating equations. Its key feature is that the two-point marginal distribution and the latent serial dependence specification can be chosen separately. The construction combines a normalized two-point distribution with a stationary latent Gaussian process via a Gaussian copula transformation, includes dependent Rademacher and Mammen multipliers, and nests the classical iid two-point wild bootstrap as the serially independent case. The induced multiplier autocovariances determine the lag weights in a corresponding heteroskedasticity- and autocorrelation-consistent (HAC) covariance estimator, which coincides exactly with the conditional covariance of the bootstrap estimating-equation sum. We establish first-order bootstrap validity for asymptotically linear estimators by showing that the matched-HAC estimator consistently estimates the long-run covariance and that the bootstrap estimating-equation sum converges conditionally to the same Gaussian limit as its original-sample counterpart, yielding valid HAC-studentized $z$-tests and the corresponding Wald and Lagrange multiplier tests. Monte Carlo experiments in nonlinear generalized method of moments (GMM) and linear regression show that Rademacher DWB generally provides more accurate finite-sample size control for $z$-tests than the Mammen and Gaussian DWB. A GMM application to a nonlinear short-rate mean-reversion model illustrates the practical relevance of the proposed two-point DWB.
Michael Coppedge, Dmitry Zaytsev, Valentina Kuskova
Tue 29 Sep
stat.MEcs.LG
pdf html
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.89
- reading political, functional, theories, functional form
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.84
Abstract
A non-significant coefficient in a dynamic panel model need not imply the absence of a relationship. It may instead reflect heterogeneous effects averaged toward zero, reciprocal dynamics overlooked by a recursive specification, or relationships masked by the omission of correlated covariates. Standard linear estimators cannot distinguish among these possibilities. We develop an inferential workflow for political panel data that resolves this ambiguity by combining flexible autoregressive estimation, forecast-necessity testing, functional characterization, and same-data linear benchmarking. The workflow first identifies relationships required for out-of-sample prediction, then characterizes their functional form across political contexts, and finally, distinguishes differences arising from estimator flexibility from those due to model specification. Applied to the causal sequence model of democratization on the V-Dem panel of 113 countries, the workflow reproduces the model's central finding - the protective belt of civil society, the rule of law, and institutionalized parties - while recovering reciprocal relationships from democracy to its institutional supports that a linear model cannot detect. Most importantly, three weak published direct effects, of which two are null, and one is marginally significant, receive three different diagnoses: one dissolves under the full specification, one reflects heterogeneous effects averaged toward zero, and one was masked by the reduced variable set. The workflow corrects the published record in both directions, removing one relationship and recovering two. More broadly, the workflow provides a framework for evaluating dynamic political theories under a model class capable of representing nonlinear and reciprocal mechanisms while preserving relationship-level interpretation and explicit inferential standards.
Jun Ma, Yuya Sasaki, Zhengfei Yu
Fri 25 Sep
econ.EM
pdf html
- interest Causal inference and identification in econometrics 0.90
- reading adjustment, gains, generic, generally
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.88
Abstract
It is standard practice to include covariates in regression discontinuity designs (RDDs) and regression kink designs (RKDs), but the theoretical justification for doing so does not generally extend beyond linear estimands. This paper proposes a novel entropy balancing reweighting approach for covariate adjustment within a general framework of RDDs and RKDs. While conventional regression-based covariate adjustment methods generally fail to deliver consistent estimation for nonlinear estimands such as quantile treatment effects, our reweighting approach achieves consistency while improving efficiency. Moreover, even in settings where the regression-based covariate adjustment method already improves efficiency, our approach can deliver additional efficiency gains. Simulation studies corroborate these theoretical findings. We present an empirical application in which our covariate adjustment yields statistically significant results that would not be obtained without covariate adjustment.
Haroon Mumtaz, Sofia Velasco
Thu 24 Sep
econ.EM
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.84
- reading shocks, macroeconomic, growth, inflation
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.84
Abstract
The tails of macroeconomic outcomes can respond differently from the centre of their distribution: shocks with modest effects on median growth or inflation can shift downside growth or upside inflation risk. We develop a threshold stochastic-volatility-in-mean VAR with regime-dependent leverage to study their structural drivers. The model allows endogenous interactions between outcomes and volatility, contemporaneous level-volatility dependence, and regime-specific propagation. In nearly 150 years of U.S. data, predictive model selection supports three inflation-defined regimes. We identify business-cycle, financial, macroeconomic-uncertainty, and financial-uncertainty shocks and decompose their contributions to growth- and inflation-at-risk. The structural composition of tail risk differs from that of the predictive median. Business-cycle shocks dominate the median response of GNP growth but account for a substantially smaller share of growth-at-risk. Macroeconomic uncertainty makes a material contribution to both growth- and inflation-at-risk, with its share of growth-at-risk increasing with the magnitude of a positive macroeconomic-uncertainty impulse, despite its limited role at the median. In high-inflation states, the contribution of financial uncertainty to inflation-at-risk rises with the magnitude of positive financial-uncertainty impulses.
Genshiro Kitagawa
Fri 25 Sep
stat.ME
pdf html
- reading kernels, series, gaussian process, processes
- interest Macroeconometrics and time series 0.91
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.85
Abstract
Gaussian processes (GPs) provide a flexible nonparametric framework for modeling time series through appropriately chosen kernel functions. This chapter introduces the basic formulation of Gaussian processes, commonly used kernels, GP regression, hyperparameter estimation, and model evaluation using in-sample and out-of-sample criteria. Applications to stationary, quasi-periodic, and seasonal time series illustrate how individual and composite kernels can represent different forms of temporal variation. Additive kernels also provide interpretable decompositions into latent components such as trend, smooth local variation, and seasonality, while product kernels allow more complex dependence structures to be constructed. Finally, Gaussian process state-space models (GP-SSMs) are briefly introduced, and a nonlinear example demonstrates how a GP transition model can be combined with particle filtering and smoothing for latent-state estimation.
Irene Botosaru, Isaac Loh, Chris Muris
Tue 29 Sep
econ.EM
pdf html
- reading equilibria, econometric, equations, counterfactual
- your work near “Valid Inference with Imperfect Synthetic Data” 0.85
- interest Semiparametric inference and debiased machine learning 0.89
Abstract
We develop a framework for identification, computation, and inference in econometric models with a linear-in-measures representation. These models express maintained restrictions as moment conditions linear in the joint probability measure of observed and latent inputs, and map that measure linearly to the distribution of outputs, even with nonlinear outcome equations. We construct an adversarial discrepancy function whose zeros characterize the identified set for structural and counterfactual parameters. With finite output support, finite linear programs compute the discrepancy function or provide certified bounds even when latent inputs have infinite support, and a penalized bootstrap yields confidence sets with uniform per-point coverage. We apply the framework to two open cases in binary choice panels with fixed effects and discrete covariates: sequential exogeneity with unspecified conditional marginal error distributions, and known conditional marginal error distributions with unrestricted serial dependence. In an entry game with multiple equilibria, the framework recovers the known sharp identification region.
Julien D. Laurendeau, Leora Sarvet, Mats J. Stensrud
Fri 25 Sep
stat.ME
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.87
- reading sequential, treatment, generalize, large
- interest Reinforcement learning, policy learning and off-policy evaluation 0.87
Abstract
Consider an optimal dynamic treatment regime, $g^{\textbf{opt}}$ correctly identified from a large, perfectly executed sequentially randomized experiment. Even when the experimental results are generalizable to a future target population, there is no guarantee that $g^{\textbf{opt}}$ outperforms human decision-makers; human experts can do better than $g^{\textbf{opt}}$ whenever they have access to relevant information beyond the covariates recorded in the experiment. Motivated by this observation, we derive results on a new class of regimes called initiation regimes, which generalize existing results on superoptimal regimes. These regimes follow human decision-makers up to the point where it becomes more beneficial to initiate a sequential optimal regime, and are guaranteed to outperform both purely human and purely algorithmic decision rules, e.g., based on reinforcement learning algorithms. Furthermore, we present modified experimental designs that identify the best initiation regimes, show how the best initiation regime can be identified from classical observational data under explicit assumptions, and give estimation and statistical inference methodology for these regimes. To illustrate the practical utility of the methods, we consider initiation regimes in a case study on treatment of lower back pain.
Stéphane Bonhomme, Koen Jochmans, Martin Weidner
Mon 28 Sep
econ.EM
pdf html
- reading nuisance, nuisance parameters, equation, likelihood models
- interest Semiparametric inference and debiased machine learning 0.92
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.85
Abstract
Many models, such as fixed-effect models for panel or network data, are hard to estimate because they feature nuisance parameters that are both numerous and estimated imprecisely. This, in general, causes an incidental-parameter problem in the estimator of the parameters of interest. The problem can be alleviated by working with an estimating equation whose expectation is insensitive to the value of the nuisance parameters. We discuss and contrast three notions of insensitivity, also called orthogonality, in the context of likelihood models: Neyman orthogonality, Neyman orthogonality to order q, and full orthogonality. Orthogonal moments are obtained by projecting the estimating equation on nested subspaces, which are spanned by, respectively, the scores of the nuisance parameters, the first q derivatives of the likelihood ratio with respect to the nuisance parameters, and all likelihood ratios of the model. We give explicit constructions in binary-choice, count-data, and nonlinear regression models.
Feras A. Saad, Todd B. Walker
Wed 30 Sep
econ.EM
pdf html
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.86
- reading economic, macroeconomic, adjustment, series
- interest Macroeconometrics and time series 0.90
Abstract
Seasonal adjustment is fundamental to economic analysis, but uncertain because seasonal components are inherently latent. This article introduces a probabilistic model discovery method that decomposes a time series into seasonal and nonseasonal components. The method returns a posterior distribution over the structure and parameters of a seasonal component. In simulation studies, the method can improve point forecasts, interval predictions, and recovery of seasonal components relative to X-13ARIMA-SEATS. In a study of eight U.S. macroeconomic series during the COVID-19 recession, the method surfaces significant ex-ante uncertainty about current seasonal adjustments in real time, well before many X-13 revisions reach their eventual peaks.
Nicolas Chopin, Andras Fulop, Yuedan Huo, Anna Simoni
Thu 24 Sep
stat.MEstat.CO
pdf html
- favorites near “Priors for the Long Run” 0.88
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.85
- reading mcmc, macroeconomics, macroeconomic, sequential
Abstract
We consider $K$-dimensional Bayesian vector autoregressions (BVARs) with Cholesky stochastic volatility (SV), in which the innovation covariance matrix is a lower-triangular linear transform of $K$ independent univariate SV processes. Such models are widely used in empirical macroeconomics to capture time-varying uncertainty and improve forecast accuracy, but the cost of posterior simulation is the binding constraint on the size of the system, and a major bottleneck for empirical work. We introduce a Markov chain Monte Carlo (MCMC) kernel that mixes better than existing samplers at the same computational complexity. It rests on a reparametrisation that makes the $K$ volatility trajectories conditionally independent, and on Particle Gibbs to update each trajectory. The kernel targets the exact posterior, rather than an approximation of it, and in an application with $K = 15$ it raises the mean effective sample size per second, relative to the benchmark corrected triangular algorithm, by a factor of approximately $14$ for the VAR coefficients and $3.4$ for the volatilities. We also introduce a a Sequential Monte Carlo squared (\smcsq{}) sampler, which used our MCMC kernel as a building block, and which delivers at every $t$ the one-step-ahead predictive density and the marginal likelihood of the data up to $t$, and hence on-line forecasting and model choice. To our knowledge, this is the first algorithm that delivers sequential marginal likelihoods for Cholesky-SV BVARs with static contemporaneous coefficients and a non-conjugate prior. We illustrate both on US monthly macroeconomic data, with $K=15$ for posterior inference and $K=6$ for the sequential sampler and model choice.
Danilo A. Sarti
Mon 28 Sep
stat.MEstat.APstat.CO
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.91
- reading economic, bayesian causal, units, causal
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.88
Abstract
Many studies expose units to several binary treatments at once and record several correlated outcomes, yet analysts usually estimate one treatment's average effect on one outcome at a time -- discarding how treatments interact and how their effects vary across units. We introduce the factorial multivariate Bayesian causal forest, a Bayesian nonparametric model that decomposes the factorial response surface over the treatment lattice into a prognostic sum-of-trees plus one sum-of-trees per main or interaction effect, each with correlated multivariate leaf parameters, estimated jointly with coherent uncertainty. A single indicator-weighted kernel samples every component -- the prognostic term is the special case whose indicator is unity -- so the two-treatment model, the single-treatment multivariate causal forest, and a general order-r truncation are one and the same sampler. We give the ANOVA/M\"obius identification of each estimand, an efficient Rcpp engine, and an interpretability layer using value-suppressing uncertainty maps. In simulations the average-effect estimators are unbiased with nominal coverage, consistent, and robust under misspecification, failing only under unmeasured confounding, which we flag. We illustrate the method on clinical, agricultural and economic data and on a deeper application to the NHANES survey. Software is provided as an R package.
Widemberg S. Nobre, David A. Stephens, Alexandra M. Schmidt, Erica E. M. Moodie
Thu 24 Sep
stat.ME
pdf html
- interest Causal inference and identification in econometrics 0.89
- reading nonparametric, target parameters, good, functional
- your work near “Valid Inference with Imperfect Synthetic Data” 0.84
Abstract
We propose a Bayesian framework for uncertainty quantification from the perspective that the working model is mis-specified in settings of a multilevel data-generating process. We focus on settings in which the mis-specification fails to match the functional form of the mean structure, and discuss Bayesian estimation of target parameters under dependence induced by a mismatch between working and data-generating models. The proposal represents a Bayesian semi-parametric procedure aimed at estimating population-level parameters while accounting for cluster- and unit-level variation in the estimating function. The proposal extends the regular Bayesian bootstrap to account for cluster- and unit-level variation using multilevel weights from an enriched Dirichlet model. Simulation studies indicate that the proposed approach has good frequentist properties when the data-generating process and the proposed model induce a partially exchangeable sequence associated with the unknown quantity of interest. Applications to radon (Gelman and Hill, 2007), Programme for International Student Assessment 2022 (OECD, 2023), and tuberculosis (Nobre et al., 2023) datasets are presented for illustrative purposes. The results demonstrate that the proposed method is competitive with variations of multilevel models, with major differences observed in the range of credible intervals, which are justified by the nonparametric assumptions underlying the proposed method.
Agnes Norris Keiller
Fri 25 Sep
econ.EM
pdf html
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.88
- interest Causal inference and identification in econometrics 0.89
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.86
Abstract
I study treatment effect estimation when treatment events have persistent effects and can be experienced more than once. Natural disasters, job loss and health shocks are examples of such treatments. I show that the effect of a total treatment trajectory can be recovered under assumptions similar to those commonly invoked in single-event settings using suitably flexible TWFE models. Decomposing the total trajectory effect into portions attributable to distinct event occurrences, however, requires further assumptions. I propose an assumption similar to conditional parallel trends, imposing it on the growth of event-specific effects rather than on untreated outcomes. Combined with a linear-in-parameters model of effect growth, this assumption enables a sequential imputation estimator that consistently estimates the dynamic effects of each event occurrence and that can accommodate heterogeneity in effects according to observable event attributes, such as intensity. I demonstrate that several intuitive TWFE models fail to recover interpretable treatment effect parameters in the multi-event setting and illustrate the sequential imputation estimator's favourable performance using Monte Carlo simulations.
Carl Bonander
Fri 25 Sep
econ.EM
pdf html
- reading aggregate, policy, units, trends
- favorites near “Difference-in-Differences with multiple time periods” 0.89
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.82
Abstract
Policy reforms are sometimes accompanied by detailed individual-level data in the implementing jurisdiction, while only aggregate outcomes are available for potential comparison jurisdictions. This article develops an identification framework for heterogeneous policy effects when individual-level data are unavailable for the comparison units. The framework combines treatment-effect contrasts from difference-in-differences comparisons within the treated jurisdiction with a compatible population-average effect identified from aggregate panel data. Identification requires relative parallel trends within the treated jurisdiction together with the assumptions needed to identify the population-average effect from the aggregate panel. The within-jurisdiction component can be estimated from repeated cross-sections with a single pretreatment period.
Arda Güçlü, Subhonmesh Bose, John R. Birge
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Eluder Dimension and the Sample Complexity of Optimistic Exploration” 0.90
- interest Adaptive data collection and anytime-valid inference 0.88
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.85
Abstract
Linear bandits model sequential decision-making problems with noisy rewards that are linear in the decision variable, where an agent must simultaneously learn about an unknown parameter that governs the mean rewards, while maximizing (expected) rewards over time. Two prominent algorithmic families--upper confidence bound (UCB) and Thompson sampling (TS)--achieve a balance of exploration (to estimate said parameter) and exploitation (utilization of knowledge about it) across time. The quality of estimation of that parameter depends on the eigenvalues of a design matrix. In this paper, we begin by showing that if the inference quality obtained from exploration, encoded in the minimum eigenvalue of the design matrix, grows $\gtrsim \sqrt{t}$ with time $t$, while actions remain sufficiently concentrated for exploitation, then an algorithm produces optimal high-probability $\mathcal{O}(\sqrt{T}\log T)$-regret rate over a time-horizon $T$ for spherical action sets. This analysis is algorithm-agnostic and follows an alternative route to the classical optimism-based elliptical-potential argument for regret analysis. Then, we illustrate that variants of UCB and TS satisfy the inference and concentration properties and in turn, enjoy optimal regret rate. In effect, our results provide a modular framework that can be used to analyze linear bandit algorithms and explicitly connect quality of parameter estimation to optimal regret accumulation.
Stefan Klus, Eirini Ioannou
Fri 25 Sep
math.DScs.LGstat.ML
pdf html
- reading equations, differential equations, operators, functional
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.91
- interest Kernels, operators and functional data 0.91
Abstract
Dynamic mode decomposition (DMD) is a data-driven method that computes the best linear approximation of the underlying dynamical system and decomposes the dynamics into a superposition of characteristic spatiotemporal patterns. Originally introduced by the fluid dynamics community, DMD and its extensions have found widespread use in many other research areas such as molecular dynamics, climate science, engineering, finance, and neuroscience. Applications include dimensionality reduction, forecasting, system identification, control, and spectral clustering. In order to apply DMD to partial differential equations, the spatial domain is typically first discretized using finite difference or finite element techniques, thus implicitly rendering the problem finite-dimensional. We extend projected and exact DMD to infinite-dimensional systems. Rather than estimating matrices from vector-valued observations, our DMD variants learn finite-rank operators from functional data such as observables, densities, or wavefunctions. We show that conventional DMD algorithms can be regarded as special cases of their functional DMD counterparts. All results will be illustrated with the aid of guiding examples. We focus in particular on Koopman, Perron-Frobenius, and Koopman-von Neumann operators associated with graphons, ordinary differential equations, and stochastic differential equations.
Sami Diaf
Tue 29 Sep
cs.LGecon.EMstat.APstat.ME
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.84
- reading monetary, economic, economics, price
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.82
Abstract
Price stability remains a pillar in monetary policy practices and carries a special importance within monetary unions. Mainstream economics tried to leverage price stability using price indices and several metrics to shed light on specific dynamics and optimal macroeconomic levels. The wide availability of data led researchers to consider the study of systems using Random Matrix Theory, based on inner correlation patterns. This aims to enhance the multivariate analysis by removing noisy patterns from the signal and improve data quality for further inferences. This work considers the collection of monthly inflation indices in the Eurozone as a \textit{system} of prices to analyze its eigenvalues' statistical and asymptotic properties and uncover inner country-level insights. Results confirm the system cannot assumed to be randomly generated, and the data exhibit noise-dominated patterns, due to small and persistent variations at the country-level. The latter make the inter-country correlations more dynamic and the separation of the signal from the noise quiet difficult. Findings identified two countries as distorting inflation dynamics besides three other distinct, regional-based groups of countries. Variability sources might stem from economic episodes fueling inflation spikes in some countries, as well as methodological aspects used to ensure data quality and representativeness in the European Union. Despite being complex, the system demonstrates a certain stability, in terms of self-organization; while large monthly fluctuations cannot be considered as rare events, but part of the data-generating process.
Jia Cai
Tue 29 Sep
stat.MLcs.LG
pdf html
- interest Semiparametric inference and debiased machine learning 0.90
- reading ard, theory, gaussian process, gp
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.88
Abstract
Automatic relevance determination (ARD) is the standard device for input selection in Gaussian process (GP) regression. By giving the covariance kernel a separate lengthscale for every input and learning those lengthscales by maximizing the marginal likelihood, ARD lets the data decide which coordinates matter: irrelevant inputs receive very large lengthscales and are effectively switched off. We trace this mechanism to the Bayesian Occam's razor embodied in the marginal likelihood, derive the gradient through which it prunes inputs, and emphasize that ARD delivers effective rather than exact sparsity. We review the algorithms used in practice and the rules that turn lengthscales into selections, and we survey the asymptotic theory, distinguishing the fixed-domain identifiability obstruction on the lengthscales from the high-dimensional selection-consistency guarantees recently established for hierarchical GP priors, and noting what remains open for plain ARD. We compare ARD with spike-and-slab priors, sparse axis-aligned and global-local shrinkage priors including the Bayesian lasso and horseshoe, penalized likelihood kriging, sensitivity and projection criteria, and additive kernels. We argue that ARD endures because of its seamless integration with kernel learning, universal software support, and low cost, and we close with its limitations and remedies.
Taewon Goo, Kihyuk Hong
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.91
- reading minimax regret, regret, argument, empirical
- interest Reinforcement learning, policy learning and off-policy evaluation 0.86
Abstract
Posterior sampling for reinforcement learning (PSRL) is one of the simplest and most effective exploration methods, but a basic question has remained open: does unmodified PSRL achieve minimax regret without structural assumptions on the prior? We answer yes. Exact vanilla PSRL is minimax optimal in leading-order Bayesian regret under arbitrary correlated priors. The difficulty is that a posterior-sampled transition model is coupled with its own continuation value. We overcome this with a common empirical transition reference that isolates the resulting value mismatch and a Bellman-based variance argument that controls it without an extra leading-order state-space factor. For finite-horizon, time-inhomogeneous tabular MDPs with unknown stochastic rewards, this yields the minimax $\widetilde{O}(\sqrt{SAH^3K})$ regret rate under arbitrary joint priors over rewards and transitions. The same proof principle gives the minimax $\widetilde{O}(d\sqrt{H^3K})$ rate for linear-mixture MDPs under arbitrary joint parameter priors.
Haijie Xu, Chen Zhang
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.91
- reading online, interventions, causal, game
- interest Adaptive data collection and anytime-valid inference 0.86
Abstract
We study fully online fixed-confidence causal discovery without any historical observational data. Starting from zero samples, the learner sequentially selects interventions to recover both the causal DAG and its edge weights under a linear-Gaussian structural equation model. We establish an instance-dependent lower bound for any $(\epsilon,\delta)$-correct algorithm and propose \textsc{FOCUS}, which adaptively allocates interventions through an online max--min game. A key contribution is a computable concentration inequality for the accumulated KL divergence involving causal parameters shared across interventions. We prove that \textsc{FOCUS} is $(\epsilon,\delta)$-correct and that its expected stopping time matches the lower bound in its $\Theta(\log(1/\delta))$ dependence up to an instance-dependent constant. Experiments demonstrate improved structure and edge-weight recovery and confirm the predicted stopping-time trend. Our codes are available on https://anonymous.4open.science/r/FOCUS_code-76E5
Hengzhi He, Guang Cheng
Thu 24 Sep
math.STstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.90
- reading log, wang, absolute, draws
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.86
Abstract
Consider testing a finite composite null $\cP={P_1,\ldots,P_L}$ against a simple alternative $Q$ using $n$ i.i.d. observations. Let $\ell_n$ denote the supremum of the expected log e-value over e-variables that are exact under every $P_i$ and pivotal across the nulls. Zhang, Ramdas, and Wang (2024) showed that $\ell_n$ is superadditive and asked whether $\ell_n/n$ converges to the upper bound $\min_i D(Q|P_i)$ under their joint atomlessness and absolute continuity assumptions. We prove that it does. The proof draws on techniques developed by Zhang, Ramdas, and Wang (2024) and by Farooq, Fritz, Haapasalo, and Tomamichel (2024).
Emmanuel Caron Parte, Jérôme Dedecker, Bertrand Michel
Fri 25 Sep
math.STstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.92
- reading short, make, nonparametric, procedures
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.88
Abstract
In this work, we study the one-dimensional regression problem under random design and Gaussian errors. Our framework is very general: we make no prior assumptions about the design (which may be nonstationary and exhibit short or long-range dependence), nor do we assume that the errors are homoscedastic. We examine in detail the cases where the error process exhibits short or long-range dependence. We adopt a least-squares penalized strategy using piecewise polynomials to estimate the regression function, following the framework of Baron, Birg{\'e} and Massart [1999]. We derive explicit penalties, up to calibration constants, to obtain adaptive estimators for which we establish risk bounds. Since these penalties depend on the dependence properties of the error process, which are unknown in practice, we propose several adaptations of the dimension jump calibration algorithm to make our procedures fully data-driven.
Javier Moral Hernández, Clara Higuera-Cabañes, Álvaro Ibraín
Mon 28 Sep
stat.MEcs.LGstat.APstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.86
- reading debt, causal, financial, dimensionality
- interest Causal inference and identification in econometrics 0.85
Abstract
This paper presents an end-to-end causal machine learning (ML) pipeline designed for real-world applications with continuous treatments. The proposed framework consists of six sequential steps: dimensionality reduction, causal identification, positivity assumption violation handling, estimation, refutation and evaluation, and policy optimization. We introduce practical contributions not currently available in existing causal ML toolkits, specifically: (1) a method for detecting and quantifying positivity violations in continuous treatment settings (2) a novel, scalable two-stage dimensionality reduction framework tailored for causal inference with high-dimensional data; (3) the adaptation of sensitivity analysis and estimation methods originally designed for binary treatments to the continuous treatment space and (4) an end-to-end integration of these components into a modular, reproducible workflow. These innovations address real-world challenges in causal inference that are often not covered in theoretical frameworks but frequently encountered in industrial applications. The methodology is validated with a synthetic dataset inspired in a real-world financial debt collection use case, however its design can be applied to analogous problems across different industries. Results demonstrate that the proposed methodology offers a more computationally efficient approach and produces less biased estimates compared to standard methods for problems with continuous treatment and high-dimensional data. A fully functional GitHub repository with documented code and numbered notebooks is made available ensuring reproducibility and practical implementation. The pipeline presented is intended to contribute to closing the gap between academic approaches and practical application in industry contexts where causal ML can be highly beneficial such as the financial sector.
Hui Wang, Edwin Fong, David T. Frazier
Mon 28 Sep
stat.MEmath.STstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.89
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.85
- favorites near “Variational Bayes under Model Misspecification” 0.91
Abstract
Martingale posteriors and related predictive resampling methods replace the likelihood--prior pair used within Bayesian inference with a predictive model for future observations. These methods are simple to implement and increasingly popular due to their computational efficiency, but little is known about their ability to accurately quantify uncertainty. In this work, we study the concentration and calibration properties of the martingale posterior for general functionals, and show that credible sets can systematically undercover if the predictive algorithms are not carefully tuned. We propose a simple remedy: the bagged martingale posterior. Rather than starting every predictive path from the observed sample, we start the paths from random bootstrap resamples of the data and then amalgamate the resulting draws. Critically, this approach incurs no additional {simulation} cost compared to standard predictive resampling, and delivers conservatively calibrated credible sets. This scalability enables application to a range of challenging examples, including sparse high-dimensional regression and nonparametric conditional quantile models.
Tianyi Ma, Kabir A. Verchand, Chao Gao, Richard J. Samworth
Thu 24 Sep
math.STstat.MEstat.TH
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.86
- interest Adaptive data collection and anytime-valid inference 0.88
- reading nonparametric, attain, implications, causal
Abstract
We consider the construction of confidence intervals for population means when observations are subject to missingness. To accommodate more general missingness mechanisms than missing completely at random (MCAR), we adopt a reparametrised version of the realisable contamination model of Ma et al. (2026), which is a mixture of an MCAR version and a missing not at random version of the same base distribution $P$. We characterise the minimax length of confidence intervals that can adapt to potentially unknown parameters of the model, including the contamination fraction, for Gaussian base distributions and for nonparametric classes satisfying certain tail or symmetry assumptions. In all of these settings, we provide explicit constructions of simple, practical and finite-sample valid adaptive confidence intervals that attain the corresponding minimax rates. Finally, we provide implications of our results for causal inference.
Chak Ming Lee, Kin Wai Chan
Tue 29 Sep
stat.ME
pdf html
- interest Causal inference and identification in econometrics 0.88
- favorites near “Difference-in-Differences with multiple time periods” 0.88
- reading nonparametric, repeated measurements, taken, achieve desired
Abstract
In this paper, we formulate a general differencing framework for variance estimation across a range of settings. We demonstrate that conventional difference-based noise variance estimators cannot achieve the desired bias-correcting power in nonparametric regression with repeated measurements. A new high-order bias-corrected differencing scheme, adapted to repeated measurements, is proposed by interlacing inter-group and intra-group differencing. The theoretical properties of the new sequences and estimators are studied. Our proposals are particularly efficient in finite samples and under high signal-to-noise ratio scenarios, where asymptotic convergence has not yet fully taken effect, due to their strong bias-correcting power.
Hassan Pazira, Marianne A. Jonker
Thu 24 Sep
stat.MEstat.CO
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.90
- interest Causal inference and identification in econometrics 0.85
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.87
Abstract
Bayesian Federated Inference (BFI) estimates statistical models from multicenter data when individual-level observations cannot be combined across centers. We present \pkg{BFI}, an \proglang{R} package that implements this methodology for Gaussian, binomial logistic, and survival regression models. Each center performs a Bayesian maximum a posteriori analysis and sends only parameter estimates and curvature information to a central server, where these summaries are combined to approximate the analysis of the combined data. The package supports prior specification, structured forms of between-center heterogeneity, several parametric and flexible baseline-hazard models for survival analysis, and treatment-effect estimation for observational and randomized studies. We describe the software design and the information exchanged between centers, give reproducible workflows for these analyses, and compare the federated results with pooled-data analyses for Gaussian, logistic, and survival models.
Hiroyuki Kasahara, Hirokazu Matsuyama, Katsumi Shimotsu, Shota Takeishi
Wed 30 Sep
econ.EMstat.ME
pdf html
- interest Causal inference and identification in econometrics 0.88
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.86
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.83
Abstract
Ignoring unobserved heterogeneity in duration models biases parameter estimates and invalidates inference, but testing for it is non-regular: the null hypothesis lies on the boundary of the parameter space and some parameters are unidentified under the null. These features render standard asymptotic theory inapplicable. This paper develops an EM test for unobserved heterogeneity in censored Weibull duration models, building on the EM approach of Li, Chen, and Marriott (2009). The test statistic has an asymptotic null distribution equal to the square of max{0, N(0,1)}, hence critical values require neither simulation nor bootstrap, and the test accommodates covariate-dependent censoring of arbitrary form. Monte Carlo simulations compare the EM test with the likelihood ratio test (LRT) of Cho and White (2010), information matrix tests, and Lagrange multiplier tests. The EM test has empirical size close to the nominal level for sample sizes of 500 or more, where the LRT remains markedly conservative and the other tests over-reject. Its size-adjusted power is comparable to that of the LRT and higher than that of the other tests. Because size adjustment requires knowledge of the data-generating process and is unavailable in practice, the EM test attains higher power than the LRT in most designs as the tests would actually be applied. In an application to the Stanford Heart Transplant data, the EM test rejects homogeneity in every covariate specification, whereas the LRT's conclusion depends on a user-chosen set of admissible parameter values and on the specification.
Shoki Okubo
Fri 25 Sep
stat.MEecon.EMstat.AP
pdf html
- interest Causal inference and identification in econometrics 0.89
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.85
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.88
Abstract
Panel conditioning, the causal effect of prior survey participation on responses, can vary with tenure. Under an additive model of cell means in period, entry cohort, and tenure, we characterize which features of the conditioning path a staggered panel identifies on its observed support, and how the unidentified component affects common panel estimators. The identified set of the path is an affine translate of the tenure projection of the cell design's kernel, and a linear functional of the path is identified exactly when it annihilates that projection. It always contains an affine direction and, when the entry cohorts share a stride, periodic directions, which exhaust it under a connectivity condition on observed increments; second differences at that stride are then identified, and ordinary ones generally are not when the stride exceeds one. Under a recruitment condition, an interrupted schedule such as the four-eight-four rotation of the Current Population Survey (CPS) distinguishes a constant increment per interview from one per calendar month, which no equally spaced schedule can. We give support conditions for recovery under a plateau, entry-wave negative controls, or bounded cohort drift. A second set of results links identification to regression: two-way fixed effects absorb every unidentified direction, so the remaining conditioning bias is normalization-invariant and itself identified, and a two-way regression with tenure indicators corrects it under a residual-rank condition. When event time is aligned with tenure, conditioning shifts event-study coefficients by a known linear functional of the path, producing pre-trends without anticipation; bounds on identified curvature bound those shifts. Simulations verify the identities, a 19-wave Japanese panel illustrates the support calculations, and published CPS month-in-sample indices give a descriptive, not identifying, example.
Yan Cui, Holger Dette, Zhou Zhou
Thu 24 Sep
stat.MEmath.STstat.TH
pdf html
- favorites near “Priors for the Long Run” 0.88
- interest Macroeconometrics and time series 0.91
- reading functional, functional time, kernels, series
Abstract
We develop a general framework for model specification testing in stationary functional time series. The approach is based on an autoregressive approximation that represents a broad class of stationary functional processes through coefficient kernels whose dimension and autoregressive order may increase with the sample size. Different model assumptions induce different structural restrictions on these kernels, and our tests are constructed by measuring deviations from the corresponding restrictions. We illustrate this principle for three problems: testing a prescribed order of a functional autoregressive model, testing a functional autoregressive moving-average specification, and testing separability of autoregressive coefficient kernels. The resulting statistics are based on weighted $\mathcal{L}^2$-distances, and the critical values are obtained by a multiplier bootstrap. We establish a quantitative bootstrap approximation that is uniform over a class of weight functions and prove asymptotic validity and consistency of the proposed tests. The methodology allows for data-adaptive weighting and is illustrated by simulations and a data example.
Wei Zhao, David M. Kaplan
Tue 29 Sep
math.STecon.EMstat.MEstat.TH
pdf html
- interest Causal inference and identification in econometrics 0.93
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.86
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.83
Abstract
Instead of having a single "yes" or "no" result from a test of the global null hypothesis that a function is increasing, we propose a multiple testing procedure of the function's increasingness at several points. If the global null is rejected, then multiple testing provides more information about why. If the global null is not rejected, then multiple testing can provide stronger evidence in favor of increasingness, by rejecting null hypotheses that the function is decreasing. Our approach uses high-level assumptions that apply to a broad class of causal and descriptive statistical models. By inverting the proposed multiple testing procedure that controls the familywise error rate, we also generate "inner" and "outer" confidence sets for the set of points at which the function is increasing. With high asymptotic probability, the inner confidence set is contained within the true set, whereas the outer confidence set contains the true set. We also improve power with stepdown and two-stage procedures. Simulation and empirical examples illustrate the new methodology, and all code is provided.
Paul Rosa
Wed 30 Sep
math.STstat.TH
pdf html
- reading semiparametric, theory, functional, linear quadratic
- interest Semiparametric inference and debiased machine learning 0.89
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.88
Abstract
We introduce a novel proof strategy for semiparametric Bernstein-von Mises theorems based on Stein's method combined with an information-geometric framework. Rather than controlling the effect of the prior on the asymptotic marginal posterior distribution of the functional of interest through the stability of an integrated likelihood under a suitable perturbation, we characterise its influence through the prior-weighted divergence of suitable vector fields over the statistical model. Applying Stein's method in this setting instead of the usual Laplace-transform approach yields explicit non-asymptotic upper bounds on the bounded-Lipschitz distance between marginal posterior distributions and the corresponding Gaussian limits predicted by semiparametric efficiency theory. These bounds consist entirely of local scalar differential quantities associated with the functional, the prior and a chosen vector field-typically related to the efficient influence function-evaluated over posterior contraction sets. We apply the theory to quadratic functionals in Gaussian white-noise models and to linear, quadratic and general integral functionals in histogram density models, with both conjugate and non-conjugate priors. For linear and quadratic functionals, the same theory identifies the differential term responsible for posterior bias and allows us to remove it through a natural functional correction, yielding Bernstein-von Mises-type theorems in regimes where they may fail for the original uncorrected functional.
Kai Ye, Shijin Gong, Hongyi Zhou, Valentina Zangirolami, Chengchun Shi
Tue 29 Sep
stat.MLcs.LGstat.ME
pdf html
- interest Reinforcement learning, policy learning and off-policy evaluation 0.89
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.85
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.87
Abstract
For decades, statistical inference has largely been developed one problem at a time. Given a scientific target, such as a treatment effect or a regression function, statisticians design a problem-specific estimator together with a procedure for quantifying its uncertainty. This paper proposes a different paradigm. We focus on a classical problem in statistical inference, confidence interval construction, and develop TabCon, an amortized inference system built on a tabular foundation model that produces confidence intervals for new datasets through a simple forward pass. The key methodological ingredients of TabCon are a sparse mixture-of-experts architecture and reinforcement-learning-based post-training that calibrate the resulting confidence intervals to a desired coverage level. Across a wide range of benchmark datasets, TabCon attains near-nominal coverage while producing short confidence intervals. At inference time, it also offers considerably greater computational efficiency, running 50 times faster than the classical bootstrap procedure, even when the latter uses only 50 bootstrap samples.
Gefei Lin, Rui Miao, Xiaoke Zhang
Wed 30 Sep
stat.MEcs.LGmath.STstat.MLstat.TH
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.86
- reading functional, theory, gains, lebesgue
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.86
Abstract
Offline reinforcement learning seeks optimal decision rules from previously collected data. In some applications, a decision can be an entire function, such as a fluence map in radiation therapy or a smooth movement trajectory in robotics. In this paper, we study the finite-sample theory for fitted Q-iteration (FQI) with functional actions in a discounted infinite-horizon setting. Three major difficulties arise in this setting: first, the absence of a Lebesgue probability density for functional actions complicates coverage descriptions; second, conventional coverage requirements can be restrictive; and third, the large functional action space makes greedy optimization in FQI challenging. To address these difficulties, we study smoothness-regularized policy search under a critic-relative coverage condition. This condition measures how well logged data distinguish relevant action-value differences without requiring an action density. Our main theorem gives finite-sample guarantees for learned-policy regret relative to the best value within a fixed smooth class of functional-action policies. The results allow trajectory lengths to be either bounded or growing and Q-functions to be fitted by either functional-input kernel ridge regression or adaptive functional neural networks. For a few examples, we can obtain polynomially decaying regret bounds in the number of logged transitions, up to logarithmic factors, with logarithmically many FQI iterations. Numerical experiments show gains of learned functional-action policies over constant-action policies and support our adoption of a critic-relative coverage condition.
Alexis-Raja Brachet, Guillaume Clavier--Frémond, Abdelhakim Ziani, Pierre-Yves Richard, Céline Hudelot
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Spectral State Space Models” 0.84
- reading forecasting, state, series, series forecasting
- interest Macroeconometrics and time series 0.87
Abstract
Time series forecasting requires extrapolating the dynamics of an observed process beyond the last available measurement. Yet recurrent forecasting models typically treat observations as inputs that directly control their latent dynamics. It leads to a regime change when these observations become unavailable at prediction time. Following a state-estimation perspective, we introduce Observer State-Space Models (OSSMs), a class of recurrent models that interprets the observed input time series as measurements of an underlying autonomous dynamical system. OSSMs explicitly separate latent-state propagation from measurement assimilation: a single transition governs the dynamics across both context and forecasting intervals, while available observations correct the estimated state through an observer. This formulation naturally exposes classical control-theoretic properties, including observability and convergence of the state estimation error. We further show that conventional and recent SSMs can be recovered as particular instances of our OSSM framework, thereby providing a unified interpretation of their recurrent dynamics and revealing modeling inconsistencies. We perform experiments across several benchmarks showing that OSSM achieves substantial improvements while maintaining the same parameter count and training setup as the corresponding SSM baseline. These results support a simple principle for recurrent forecasting: observations should correct the estimated latent state, rather than control the dynamics used to propagate it.
Maoyu Zhang, Jingfei Zhang, Xuming He
Tue 29 Sep
stat.ME
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.86
- reading gains, inference, prediction-powered inference, predictive distributions
- interest Online learning and statistical learning theory 0.87
Abstract
Prediction-powered inference (PPI) typically relies on point predictions on unlabeled data. When predictive distributions are available as in a wide range of applications, including predictions from LLMs, we introduce distribution-informed prediction-powered inference (DiPPI), a general framework for further improving statistical efficiency by using predictive distributions as auxiliary information in the spirit of PPI. We characterize the optimal use of this information through score calibration, derive the oracle efficiency for a finite-dimensional representation of the predictive distribution, and provide theoretical guarantees for positive learning with the cross-fitted DiPPI estimator. Through simulations and three real-data applications, we show that DiPPI achieves better efficiency than PPI methods based on point predictions. These gains arise when the predictive distribution contains score-relevant information that is partially lost in point predictions.
Michael Pfarrhofer, Anna Stelzer
Wed 30 Sep
econ.EMstat.AP
pdf html
- reading monetary, monetary policy, policy, shocks
- favorites near “SVAR (Mis)Identification and the Real Effects of Monetary Policy Shocks” 0.88
- interest Macroeconometrics and time series 0.85
Abstract
We answer the question posed in the title with a nonlinear mixed-frequency vector autoregression, estimated with Bayesian additive regression trees. The model combines monthly macro-financial variables with quarterly bank lending survey data, and identifies the dynamic responses from high-frequency policy surprises. Sign asymmetry dominates; a tightening produces monetary policy transmission mostly in line with the theoretical predictions, whereas easing of any size produces mostly insignificant responses. Peak effects vary with initial conditions, while several common, predefined regime splits do not yield considerable differences in the propagation of the shocks.
Aidan Wardak, Sayar Karmakar ( University of Florida)
Mon 28 Sep
econ.EM
pdf html
- favorites near “Priors for the Long Run” 0.86
- interest Macroeconometrics and time series 0.89
- reading white, nearly, procedures, large
Abstract
We derive the limiting distributions of the $M$-test family of unit root statistics in the nearly integrated nearly white noise (NINW) framework introduced by Nabeya and Perron (1994) in the case of an unknown linear time trend. In the case of known long run variance (LRV), the limiting distributions of the $M^{GLS}$ tests are contaminated by additional noise terms as a result of quasi differencing whereas these terms are less present in the $M^{OLS}$ limiting distributions, both of which display conservative properties under conventional critical values. Furthermore, we prove the Gaussian power envelope in the NINW model is asymptotically equivalent to the standard envelope of Elliott, Rothenberg, and Stock (1996), and that the oracle $M$-tests have inefficient power relative to this benchmark. We then derive the limiting distributions of the feasible statistics and show that the autoregressive estimate of the LRV commonly used overestimates the LRV, creating altered limiting distributions. Finally, finite sample simulations illustrate that, of the procedures considered, no uniformly satisfactory solution exists for handling a series with a large negative moving average coefficient.
Bruno Ebner, Dominic Edelmann, Norbert Henze, Frédéric Ouimet, Donald Richards
Fri 25 Sep
math.STmath.CAmath.FAmath.PRstat.TH
pdf html
- reading operator, operators, eigenvalues, equations
- interest Semiparametric inference and debiased machine learning 0.88
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.87
Abstract
The Baringhaus-Henze-Epps-Pulley (BHEP) tests for multivariate normality are affine-invariant goodness-of-fit tests based on a Gaussian-weighted $L^2$ distance between empirical and Gaussian characteristic functions. In 1990, Henze and Zirkler expressed the limiting null distribution through the eigenvalues of an integral operator on the standard Gaussian space. In 1997, Henze and Wagner obtained a simpler covariance kernel and raised the problem of calculating the eigenvalues of the resulting operator on a Gaussian-weighted space. Although subsequent work treated the univariate case and numerical approximations in a few low dimensions, the complete all-dimensional spectral problem remained open. This paper determines both complete spectra for every dimension $d \in \mathbb{N}$ and every smoothing parameter $\beta > 0$. The two operators are shown to have the forms $\mathcal{X}_{\beta,d}^*\mathcal{X}_{\beta,d}$ and $\mathcal{X}_{\beta,d}\mathcal{X}_{\beta,d}^*$ for the same Hilbert-Schmidt operator $\mathcal{X}_{\beta,d}$. Consequently, their nonzero eigenvalues agree, including multiplicities, while the null space of the Henze-Zirkler operator is identified exactly. The Gaussian integral operator in the Henze-Wagner decomposition is diagonalized by Mehler's formula, and rotational symmetry confines the finite-rank correction to the sectors associated with spherical harmonics of degrees $0$, $1$, and $2$. The degree-$1$ and degree-$2$ eigenvalues are characterized by scalar transcendental equations, and the radial eigenvalues by an explicit pole-safe Fredholm determinant. The paper establishes nonnegativity, multiplicities, eigenfunction reconstruction, completeness, the trace identity, and a complete characterization of all exceptional pole cases.
Jia Cai
Tue 29 Sep
stat.MLcs.LGcs.NAmath.NA
pdf html
- interest Semiparametric inference and debiased machine learning 0.90
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.87
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
Abstract
Automatic relevance determination (ARD), the default tool for variable selection in Gaussian-process (GP) regression, ranks inputs by inverse lengthscales -- which measure how fast a function varies, not how much an input contributes to prediction -- and offers no calibrated rule for deciding which inputs to keep. The prediction-centred alternative, the derivative sensitivity $\nu_j = \mathbb{E}[(\partial f/\partial x_j)^2]$, is available in closed form from a fitted GP, but turning it into a selection rule is harder than it looks: at a null input the estimator is a degenerate quadratic form, so Wald and Bernstein-von Mises cutoffs are anti-conservative, and the natural residual bootstrap is mis-scaled. We show that a studentized multiplier bootstrap of the GP derivative process repairs both, prove its validity through an invariance principle for quadratic forms, and obtain asymptotic family-wise and false-discovery-rate control across inputs. Over 100 replications the rule controls FDR wherever inputs are truly null, while uncalibrated derivative rankings breach the target by up to 2x and a Bernstein-von Mises cutoff by 2.2x; at matched FDR it loses no power; it holds under a Mat\'ern kernel and input correlation up to 0.99; on real data with planted and authentic null inputs it admits 5-12x fewer spurious inputs; it costs 5-18% of the GP fit; and a block-averaged variant retains validity at cost linear in $n$.
Ariel Linden
Thu 24 Sep
stat.ME
pdf html
- interest Causal inference and identification in econometrics 0.90
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.87
- favorites near “Difference-in-Differences with multiple time periods” 0.86
Abstract
Rubin's variance ratio (VR) complements the standardized mean difference by detecting covariate imbalance in spread, but no purpose-built omnibus extension exists for more than two groups. We introduce FVR, a size-weighted quadratic combination of pairwise log-variance-ratios that generalizes VR to k groups, together with geometric-mean and maximum pairwise variance-ratio comparators. FVR has an exact relationship to Rubin's VR at k = 2 and a mathematical structure directly parallel to Cohen's f. In Monte Carlo simulations spanning four variance-driven bias mechanisms, k = 3, 4, 6, and sample sizes from 200 to 10,000, FVR and the geometric-mean statistic generally tracked downstream estimation bias at least as well as the maximum statistic. FVR retained substantially more signal when many groups shared a very small sample, while absolute-bias thresholds were unstable at N = 200. Across mechanisms, FVR < 0.10 was reasonably reassuring, values above approximately 0.30 generally indicated concern, and intermediate values were context-dependent. The statistic is implemented for arbitrary analysis weights in the Stata command varatio and illustrated using a multivalued-treatment application.
Isfandiyor Akhmedov
Fri 25 Sep
stat.ME
pdf html
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.86
- interest Causal inference and identification in econometrics 0.88
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.83
Abstract
Competing-risks data do not identify latent marginal duration distributions or their dependence without additional restrictions. This paper asks whether restrictions introduced to restore identification themselves restrict the observable law. For a two-risk Archimedean model with categorical exclusion restrictions, we derive a necessary-and-sufficient observable characterization. A discrete single-crossing argument identifies the scalar copula parameter from cell-specific overall survival probabilities, while cause indicators recover the remaining allocation and generate additional specification restrictions. We construct an identification-robust, self-normalized quadratic statistic, invert it to obtain confidence sets, and use empty inverted sets as a conservative specification test. Simulations show that the cause indicator can be decisive: under the weakest contrast considered it converts a frequently uninformative survival-based confidence set into an informative joint set without loss of coverage, whereas excessive categorical contrast can eliminate causes from individual cells and make cause-specific recovery inadmissible. The results turn latent exclusion restrictions into refutable restrictions without estimating covariate derivatives.
Mohammad Fesanghary
Mon 28 Sep
cs.LGstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.85
- interest Causal inference and identification in econometrics 0.85
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.84
Abstract
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Mar\v{c}enko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum indicates pervasive factor confounding, LUCID attenuates factor-dominated variation and recovers contemporaneous (lag-$0$) structure from the resulting innovations, with edge selection calibrated against a data-driven edge-free null. Rather than being tied to a particular discovery algorithm, it can wrap existing discovery engines; we demonstrate consistent improvements across three such methods. On a diverse synthetic out-of-distribution benchmark spanning changes in confounder strength and sparsity, loading density, lag structure, volatility dynamics, edge heterogeneity, persistence, intermittency, and tail behavior, LUCID achieves the best family-weighted directed, lag-resolved graph $F_1$ ($0.60$), improving over the strongest baseline by $0.19$ absolute ($\approx\!46\%$ relative). Its advantage widens relative to looser lag-collapsed scoring, and remains robust under intermittent and heavy-tailed confounding. Code reproducing the method, the benchmark generators, and every reported experiment is available at https://github.com/bloomberg/causal-ts.
Ulrich Hounyo
Thu 24 Sep
econ.EMstat.APstat.ME
pdf html
- favorites near “Difference-in-Differences with multiple time periods” 0.89
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.85
- interest Causal inference and identification in econometrics 0.87
Abstract
Staggered distributional difference-in-differences produces cohort-specific potential-outcome distributions, but applied work typically wants one overall quantile treatment effect. Two natural summaries---averaging cohort quantile treatment effects (QTTs) and mixing cohort distributions before inversion---use the same policy weights yet answer different target-population questions and can disagree even in sign. We derive the exact sharp interval for their gap conditional on cohort quantiles and weights, a globally sharp range-only envelope, and a locally sharp density-tilt representation. We develop joint smooth and mass-point-safe inference and show that neither estimator is uniformly more precise, even when the estimands coincide. In a same-object reconstruction of a public staggered-QTT application, holding data, identification, distributions, and weights fixed while changing only aggregation order reverses reported signs at several quantiles. Aggregation order is therefore part of the estimand and must be chosen before inversion.
Hongye Yang, Boxiao Huang
Fri 25 Sep
cs.AIstat.ML
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.84
- interest Reinforcement learning, policy learning and off-policy evaluation 0.85
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.86
Abstract
Tool-using agents are continually updated with new interaction data. After each policy update, however, previously estimated action credits may become stale. Recomputing them from scratch can require many additional tool calls and environment interactions, making repeated updates increasingly expensive. We ask a simple question: when does historical action credit actually need to be updated? Our key observation is that a change in action value does not necessarily imply a change in the decision. Historical credit can still be useful as long as policy-induced drift is too small to overturn the existing action ranking. Building on this idea, we introduce pairwise branch sensitivity to capture how strongly a policy update affects the downstream regions that distinguish two candidate actions. We then derive a first-order anchored credit-transport estimator that updates historical credit using old interventional trajectories, and propose a Decision-Sufficient Credit Gate (DSC-Gate) that chooses whether to reuse, transport, or resample credit. Experiments show that branch sensitivity explains credit drift substantially better than global policy distance. With sufficient historical data, credit transport reduces estimation error, while its benefit to decision making is concentrated on updates that affect action-distinguishing branches. On a fully independent test set, DSC-Gate changes mean regret by only +0.00004 relative to a gap-based gate while reducing mean new tool steps from 472 to 286, a 39.4% reduction. We observe the same pattern after a real tool-agent parameter update. Overall, our results show that agents do not need to recompute action credit after every policy update: much of the historical evidence can be reused or cheaply corrected, reducing the additional interaction required to keep action decisions up to date.
Tao Jiang, Lin Xiao
Wed 30 Sep
math.OCcs.AIstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.85
- reading equations, conceptual, special, moment
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.83
Abstract
Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first stage focuses on the analysis of conceptual (impractical) methods that rely on the exact first and second moments. In the second stage, we replace the exact moments with their respective estimators, and invoke Dvoretzky's theorem to show that the resulting practical methods converge almost surely to a neighborhood of the target solution. The size of the neighborhood depends on the biases and variances of the first- and second-moment estimators. We derive concrete bounds for Muon and a spectral variant of Adam that determine the radius of their neighborhood of convergence.
Yang Cai, Vineet Gupta, Yanchen Jiang, Christopher Liaw, Aranyak Mehta, Grigoris Velegkas, et al.
Thu 24 Sep
stat.MLcs.LG
pdf html
- reading cumulative regret, regret, online learning, tuned
- interest Online learning and statistical learning theory 0.90
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.83
Abstract
Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate tuned to $T$, and is known to be tight. If instead the regret bound is required to hold simultaneously at every time $t$, the best known guarantee has been $\sqrt{t \ln n}$---a factor of $\sqrt{2}$ worse---and it has remained unknown whether this factor of $\sqrt{2}$ is necessary. We show that it is not. We give an algorithm, requiring no knowledge of the horizon, whose cumulative regret satisfies $R_t \le \bigl(1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ simultaneously for every $t \ge 1$.
Qilan Hong, Runze Li
Tue 29 Sep
stat.AP
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.81
- reading technology, financial, market, investment
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.78
Abstract
This paper conducts an empirical mediation analysis to examine how U.S. stock-market sectors responded to the 2025 tariff-policy episode. Using S&P 500 stock data along with hundreds of firm-level financial variables, we study whether sector membership was associated with differences in stock returns and whether these differences were statistically related to firm financial characteristics. For each of five economically distinct tariff-policy windows, we use high-dimensional linear causal mediation models to estimate the direct sector effects. We use a linear regression model to estimate the total sector effects, and then estimate the indrect sector effects by using the difference between the total sector effects and the direct sector effects. To accommodate a large number of candidate mediators, we use a partially penalized least-squares procedure that regularizes financial-variable coefficients while unpenalizes coefficients of all sector indicators. Wald and F-type tests are used to examine whether indirect and direct effects are significant or not, respectively. The results show that tariff news did not produce one uniform market response. Sector effects vary across five economically distinct tariff-policy windows. The clearest findings are in the direct sector effects. Technology was the sector most negatively affected during the initial decline and escalation collapse, but it also rebounded most strongly during the April 9 relief rally. The results indicate that in addition to Technology, Energy and other trade-exposed sectors were particularly affected during the initial decline and escalation phases. The selected mediators provide evidence that growth, investment, liquidity, profitability, and cash-flow characteristics helped characterize firms associated with sector-level return differences.
Zijun Chen, Zihan Zhang
Wed 30 Sep
cs.LGstat.ML
pdf html
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.89
- your work near “Timing as an Action: Learning When to Observe and Act” 0.84
- interest Reinforcement learning, policy learning and off-policy evaluation 0.83
Abstract
We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $\epsilon$-optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error $V_{0}^{*, m} - V_{0}^{\widehat\pi^{m}, m}$ where $m\in [M]$ represents the reward function and $V_{0}^{*, m}=\mathbb{E}_{s_1\sim \mu}[V_{1}^{*, m}(s_1)]$. Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of $$ O\left(\frac{SAH^3}{\epsilon^2}\log M \mathrm{polylog}\left(\frac{SAH\log M}{\min\left\{\epsilon, 1\right\}\delta}\right)\right)$$ episodes, with no additional burn-in cost. This matches the information-theoretic lower bound up to a factor of $ \mathrm{polylog}(SAH\log M/(\min\left\{\epsilon, 1\right\}\delta))$. Our method combines three technical ingredients. First, we adapt MVP to reward-switching learning to construct optimistic value estimates. Second, we use fresh replay samples to conservatively evaluate the candidate policies. Third, gap-based multiplicative weights updates adjust the reward-sampling distribution using the differences between these estimates, converting weighted learning progress into simultaneous guarantees for all rewards.
Daniel Barzilai, Ohad Shamir
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.86
- interest Online learning and statistical learning theory 0.87
- reading generalize, special, modern, interpolation
Abstract
Understanding why predictors can generalize despite interpolating noisy training data is a central puzzle in machine learning. Most work on such "benign overfitting" studies minimum-2-norm linear regression, reflecting the inductive bias of gradient descent. However, modern optimizers such as Adam and Muon use non-Euclidean update geometries, favoring solutions associated with other norms. Analyzing regression for non-Euclidean norms is substantially more difficult, with known results essentially limited to Gaussians. In this paper, we develop a method to analyze benign overfitting in linear regression for general norms and general (sub-Gaussian) distributions. As a special case, we prove that minimum-p-norm interpolation with p>1 can benignly overfit even for non-Gaussian distributions, under suitable conditions. Perhaps surprisingly, for the 1-norm, benign overfitting does not hold in general for well-behaved (but non-Gaussian) distributions, showing that existing positive 1-norm results rely crucially on Gaussianity. Our proof analyzes the geometry of the dual optimization problem, using concentration and central limit tools to show it is approximately Euclidean in many high-dimensional cases.
Ruizi Yan, Dennis Frauen, Maresa Schröder, Stefan Feuerriegel
Fri 25 Sep
stat.MLcs.LG
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.89
- reading causal, harm, indirect, goal
- interest Causal inference and identification in econometrics 0.83
Abstract
A central goal when designing treatment policies is often to "do no harm", that is, to avoid interventions that improve average outcomes while worsening outcomes for some individuals. A widely used notion for harm is the fraction of negatively affected (FNA), defined as the probability that an intervention decreases an individual's outcome. However, in many applications, treatments operate through mediators, and a single "total" FNA can obscure whether harm arises primarily through direct pathways or indirect (mediator-induced) pathways. In this work, we introduce a path-specific analogue of the FNA. For this, we disentangle total harm into direct and indirect harm in causal mediation settings. However, these quantities depend on joint distributions of potential outcomes that are not point-identified even in randomised controlled trials. As a remedy, we develop a novel partial identification framework for direct and indirect FNA. In our framework, we (i) derive sharp Makarov bounds for the FNA, and (ii) propose a semiparametrically efficient estimator with valid confidence intervals for these bounds under mild margin conditions. We demonstrate our framework across various numerical experiments. To the best of our knowledge, we are the first to study path-specific decomposition of causal harm and to develop an orthogonal inference framework for its analysis.
Md. Niamul Islam Sium, Shafayet Khan Shafee, Mohammad Hridoy Patwary, Bishal Sarker
Tue 29 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.88
- interest Causal inference and identification in econometrics 0.85
- reading indirect, absolute, degree, consistent
Abstract
The proportion mediated (PM), defined as the indirect effect divided by the total effect, is intuitive when direct and indirect effects operate in the same direction. Under inconsistent mediation, however, opposing pathways partially cancel each other, causing PM to fall outside the unit interval or become unstable as the total effect approaches zero. We introduce two complementary effect-scale measures: the Mediated Magnitude Share (MMS), which quantifies the absolute magnitude of an indirect pathway relative to the combined absolute magnitudes of the direct and indirect pathways, and the degree of suppression, which quantifies the fraction of this combined magnitude that is cancelled by opposing pathways. We establish the bounds and limiting behavior of both measures in the two-pathway case, show that MMS reduces to PM under consistent mediation, and extend both measures to settings with multiple indirect pathways. For binary outcomes, we extend the framework to odds-ratio and risk-ratio scales by using logarithms to convert multiplicative effects into additive ones, enabling direct application of MMS and the degree of suppression. Together, these measures provide bounded and interpretable characterizations of mediation systems in which pathway directions are not uniform.
Hang Xu, Yi Shen, Yuzhong Zhao
Mon 28 Sep
math.STcs.ITmath.ITstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.87
- favorites near “Optimal uniform convergence rates and asymptotic normality for series estimators under weak dependence and weak conditions” 0.90
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.82
Abstract
We investigate the least squares linear regression problem with random partial Discrete Fourier Transform (DFT) matrices, providing a rigorous analysis of the model's generalization error. By leveraging tools from random matrix theory, we derive exact non-asymptotic bounds for the risk of the Moore-Penrose estimator, which hold for finite-dimensional problems and reveal the precise dependence on key parameters such as the sample size, dimension, and noise variance. Then we obtain a characterization of the double descent phenomenon in the linear regression context, demonstrating how the risk evolves when the number of parameters $p$ and the number of samples $n$ tend to infinity, with $p/n$ fixed. The analysis relies on applications of the Stieltjes transform for random Fourier matrices, enabling a precise description of the spectral properties of these matrices and their impact on regression performance. To validate our theoretical findings, we present several numerical examples that illustrate the double descent curves. These simulations align closely with our derived bounds, confirming their predictive power in both under-parameterized and over-parameterized regimes.
Lars Kühmichel, Stefan T. Radev, Bhanu Prasanna Koppolu, Masoumeh Davoudi, Jerry M. Huang, Paul-Christian Bürkner
Wed 30 Sep
cs.LGstat.ML
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.85
- interest Computational methods for heterogeneous agent and dynamic economic models 0.81
- favorites near “Spectral State Space Models” 0.79
Abstract
Neural simulation-based inference (SBI) has been widely successful in inferring a relatively small number of interpretable parameters from potentially high-dimensional observations, such as images or time series. Accordingly, representation learning in SBI has focused almost exclusively on compressing the observations used to condition the posterior. More recently, however, SBI has begun to target increasingly high-dimensional parameter spaces, raising the complementary question of whether the inference target itself should be compressed. Our answer is a practical merger of SBI and latent generative modeling, which learns a low-dimensional representation of the simulator parameters, performs posterior inference directly in this latent space, and maps posterior samples back to the original parameter space. We characterize the conditions under which latent-space inference recovers the desired target posterior and systematically study its empirical trade-offs. Across four case studies and three generative families, we compare latent and standard estimators while controlling for network capacity, regularization, optimization, and training compute. At matched training compute, latent-space inference achieves accuracy and marginal calibration comparable to direct target-space inference while sampling up to more than an order of magnitude faster.
Subir Hait
Fri 25 Sep
stat.ME
pdf html
- interest Semiparametric inference and debiased machine learning 0.90
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.89
- favorites near “Difference-in-Differences with multiple time periods” 0.84
Abstract
An analyst should not assume that a design score improves a meta-analysis simply because it is labeled as robustness. This paper treats declining heterogeneity along a prespecified, outcome-separated design score as a falsifiable directional hypothesis. In an 18-study illustration, the constrained model lands at $\gamma = 0$ and becomes numerically identical to conventional random effects, while the unrestricted scale fit points in the opposite direction. Simulations that deliberately violate the restriction show the same graceful fallback under wrong-direction and U-shaped heterogeneity; a stepwise decreasing variance function retains an efficiency gain despite functional-form misspecification. DR-Meta formalizes the hypothesis as $\tau^2(\mathrm{DR}) = \tau_0^2 \exp(-\gamma \mathrm{DR})$, $\gamma \geq 0$, within an established location-scale framework. Exact nesting, conditional weight monotonicity, a fixed-effect upper bound, pseudo-true targets under mean misspecification, and score-rescaling behavior are derived with full proofs. Efficiency gains are negligible around an empirically anchored $\gamma$ near 1 and become material only under stronger gradients; the largest RMSE reduction, about 9.6 percent, occurs in a deliberately extreme condition. Plug-in intervals under-cover relative to random effects, and modified Knapp-Hartung adjustment improves coverage only modestly. Widening the $\gamma$ optimization bound changes $\gamma$ estimates much more than pooled-estimate RMSE. The practical contribution is therefore diagnostic rather than automatic efficiency: a nested directional restriction can be supported, contradicted, or collapse transparently to random effects, provided constrained, unrestricted, location, and sensitivity analyses are reported together.
Junghyun Lee, Minsoo Ha, Sanghwa Kim, Yeongjong Kim, Eunjee Lee, Seiyun Shin, et al.
Wed 30 Sep
cs.LGstat.ML
pdf html
- interest Reinforcement learning, policy learning and off-policy evaluation 0.88
- your work near “Timing as an Action: Learning When to Observe and Act” 0.84
- favorites near “Preference-based Online Learning with Dueling Bandits: A Survey” 0.87
Abstract
Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairwise losses outperform pointwise losses? We study this question in a grouped offline contextual-bandit setting allowing multiple actions per context, capturing many reward learning scenarios. We compare Value Regression (VR), which regresses observed rewards pointwise, with Value Difference Regression (VDR), which regresses reward differences between a pair of actions sampled under the same context. We consider a semiparametric model where the mean reward is the sum of a learnable action-dependent component and an arbitrary context-dependent yet action-independent nuisance, capturing context-specific disturbances. Using a unified localized analysis, we prove finite-sample regression guarantees for finite and linear function classes and translate them into offline-regret bounds. For finite classes, VDR eliminates the misspecification term in the VR bound and improves a reward-scale-dependent error term by averaging over actions within each context, a benefit absent from the corresponding VR term. For linear classes, neither method uniformly dominates: within-context differencing removes nuisance-induced bias but may increase estimation variance relative to using absolute rewards when the misspecification is sufficiently low. This yields a feature geometry-dependent bias-variance tradeoff, which we corroborate with numerical experiments.
Xingwei Liu, Yuhong Yang, Wangli Xu
Wed 30 Sep
stat.MEmath.STstat.MLstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.91
- your work near “Valid Inference with Imperfect Synthetic Data” 0.82
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.86
Abstract
Goodness-of-fit testing is a basic tool for assessing whether a fitted procedure has captured the systematic information contained in the covariates. While traditional theory has largely focused on parametric regression models, modern data analysis increasingly relies on flexible black-box learners, whose predictive success alone is insufficient to assess model accuracy. In this paper, we propose SPARK, a general framework for goodness-of-fit testing that applies to traditional statistical models and general black-box learning procedures, continuous and binary responses, and low- and high-dimensional predictors. Based on a debiasing strategy, the residuals from an initial fit of a learning procedure are projected onto nearly orthogonal directions to extract any remaining signal. To capture information across all projection directions, we propose a kernel-based projection method and establish both its asymptotic properties and the consistency of a bootstrap procedure. Comprehensive simulations and real data analyses illustrate the effectiveness and flexibility of our proposed method.
Fernando Pires Hartwig, Neil Martin Davies, George Davey Smith
Mon 28 Sep
stat.ME
pdf html
- reading causal, standard errors, estimating causal, effect exposure
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.89
- interest Causal inference and identification in econometrics 0.83
Abstract
Two-sample Mendelian randomization (MR) is a widely applied methodology in epidemiology. In two-sample MR, summary data (typically, regression coefficients and standard errors) quantifying the association between multiple genetic variants and the exposure and the outcome are used in an instrumental variable framework aimed at estimating the causal effect of the exposure on the outcome. Most two-sample MR methods were developed under data-generating models where the association of for each candidate genetic instrument with the exposure, as well as the causal effect of the exposure on the outcome, are constant in the additive scale. These assumptions are useful because they imply that, had all genetic variants been valid IVs, they would all estimate the same causal parameter - namely, the constant causal effect. We refer to this condition as summary-level homogeneity. However, these are rather strong homogeneity conditions which may raise concerns about the plausibility of these methods in practice. In this paper, we show that summary-level homogeneity is implied by the following conditions: the causal effect is additive linear, but not necessarily constant across, all strata of the population; and uncorrelatedness between heterogeneity in the causal effect and in the association between each genetic variant and the exposure. Under these conditions, typical two-sample MR methods can be interpreted as estimators of the average causal effect. These results clarify that point-identifying assumptions required for two-sample MR methods are weaker than previously anticipated, which contributes to their plausibility and interpretation in at least some practical applications.
Bruno Ebner, Norbert Henze
Tue 29 Sep
math.STstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.89
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.89
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.79
Abstract
The Baringhaus--Henze--Epps--Pulley (BHEP) tests form a widely applicable class of affine invariant and consistent tests for multivariate normality. Under the null hypothesis, the test statistic converges to a weighted sum of independent chi-squared random variables. Although closed-form expressions for the first three cumulants of this limiting distribution are known for arbitrary dimension $d$ and smoothing parameter $\beta$, an explicit expression for the fourth cumulant has so far been available only in a special univariate case. We derive the fourth cumulant for arbitrary $d\geq1$ and $\beta>0$ and use the resulting first four moments to construct Johnson and Pearson approximations to the limiting null distribution. These yield essentially instantaneous analytic approximations to BHEP $p$-values, without numerical eigenvalue calculations. Using the recently obtained complete spectrum as a benchmark, we show that both four-moment approximations are highly accurate over a broad range of dimensions, smoothing parameters and upper-tail probabilities, and substantially improve on two- and three-parameter lognormal approximations. Monte Carlo results indicate good finite-sample calibration for many parameter combinations, although convergence to the limiting distribution may be slow in higher dimensions for extreme values of the smoothing parameter.
Mehrdad Pournaderi
Wed 30 Sep
stat.MEecon.EMstat.AP
pdf html
- favorites near “Priors for the Long Run” 0.86
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.83
- interest Causal inference and identification in econometrics 0.85
Abstract
We test whether a multivariate vector X conforms to a specified distribution F, a problem in copula modelling and density forecasting. The Rosenblatt transform reduces it to a test of uniformity, but depends on an arbitrary coordinate ordering that strongly affects power under dependence. We study order randomization: applying the transform under many random orderings and merging the evidence with dependence-robust rules. Reordering conserves the total Mahalanobis signal energy and merely redistributes it, so one ordering is a lucky or unlucky draw. In simulations we observe significant gains in calibrated power over both the expected single random ordering and order-invariant references. Two ingredients are essential: a two-sided base statistic, and a re-estimating parametric bootstrap that restores level under an estimated null and unlocks the gain. The calibrated pooled tests are robust to the departure's shape; no order-invariant reference we compare is: the symmetric-root test collapses on diffuse departures, while the shape-flat chi-squared test trails on concentrated ones. We apply it to a Gaussian foreign-exchange risk model over a decade of daily data on nine currencies, where it detects episodes such as Brexit and COVID. Though we focus on Gaussian nulls, the procedure extends to any null whose conditional distributions can be computed and simulated from.
Jakub Marecek, Enrico Biffis, Abigail Langbridge, Robert Shorten
Tue 29 Sep
math.STcs.LGstat.TH
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.85
- interest Computational methods for heterogeneous agent and dynamic economic models 0.82
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.82
Abstract
A special semimartingale admits a unique decomposition $X=X_0+M+A$ into a local martingale $M$ and a predictable finite-variation part $A$. We consider the identification of $A$ when $X$ is observed only through repeated cross-sections. The estimand is then the projection of the sampled predictable compensator onto the observable feature filtration, namely the current state together with whatever randomness is shared across the population, so that at a fixed diffusion coefficient the marginal flow identifies the drift only up to a Markovian projection. If the drift is an affine functional of an observed lag window, the joint problem is a convex quadratic programme whose solution is the pseudo-panel regression of econometrics. Our principal concern is the case, which we believe not to have been treated before, in which the drift is the output of a hidden linear dynamical system whose dynamics are themselves to be identified from the marginals. The joint problem is then a bilinear quadratically constrained programme, which we solve to certified global optimality by spatial branch and bound; with unpenalised state disturbances and a drift basis growing with the grid it is NP-hard already in latent dimension one, by reduction from $\ell^1$ rank-one matrix approximation, whereas the complexity of the deterministic system at fixed latent dimension remains open. A block-coordinate decomposition offers a cheaper alternative. For the estimator itself, we obtain rates at a fixed mesh, separated into Monte-Carlo, estimation and grid contributions.
Athénaïs Gautier
Tue 29 Sep
stat.MEstat.AP
pdf html
- reading gaussian process, mcmc, nonparametric, quadrature
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.76
- interest Semiparametric inference and debiased machine learning 0.85
Abstract
The SLGP package provides nonparametric estimation of fields of probability densities indexed by covariates, from unevenly sized samples with few or no replicates. A Spatial Logistic Gaussian Process applies a logistic density transformation to a finite-rank Gaussian process over a joint index-response domain, yielding a prior over conditional densities that adapts in location, shape, and modality. The contribution of this paper is the implementation: a finite-rank Random-Fourier-Feature representation that keeps the transformed process numerically tame, a log-concave likelihood that makes the fit a convex problem, and a family of quadrature shortcuts for the normalising integrals that turn a per-observation cost into a per-grid cost. Three estimation modes are available: MAP, Laplace, and full MCMC. We demonstrate the package throughout on the classical Fiji-Tonga earthquake catalogue, estimating how the distribution of hypocentre depths deforms across the region, and on a discrete-response variant.
Howell Tong
Tue 29 Sep
stat.ME
pdf html
- reading econometric, econometrics, references, author
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.84
- interest Macroeconometrics and time series 0.88
Abstract
This is mostly a non-technical note, reflecting the author's philosophy and personal views on time series analysis, with references limited to only a few representatives for brevity. The Threshold Principle, formally announced in Tong (1990), modernised time series analysis by introducing a collection of sub-systems to model complex nonlinear dynamics. We explore the conceptual architecture of the Threshold Principle drawing parallels and contrasts with artificial neural network in machine learning. We trace the influential adoption of Threshold Autoregression in econometrics. We examine a smooth extension, namely the smooth threshold autoregression introduced by Chan and Tong (1986) that was later popularized in the Econometric literature. We sound cautions to help econometric users to avoid misuse of this model. Furthermore, we examine how we can systematically apply the Threshold Principle to conditional variance to enable meaningful volatility classification. Finally, we mention some of the modern applications of the Threshold Principle to non-real-valued domains, underscoring its enduring half-century methodological significance as embodied in the threshold autoregression.
Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong
Tue 29 Sep
cs.LGstat.ML
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.85
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.89
- interest Reinforcement learning, policy learning and off-policy evaluation 0.84
Abstract
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficiency. To address these issues, this paper introduces a weighted Bellman residual minimization framework that incorporates density ratio weighting by effectively integrating expert demonstrations with behavioral data. The proposed weighting scheme departs from the conventional completeness assumption commonly imposed in the theoretical analysis of deep reinforcement learning. We establish a sharp convergence rate for density ratio estimation and derive the convergence rate for the excess risk of resulting deep $Q^*$ estimator. Extensive empirical evaluations demonstrate that, compared to existing methods, our method achieves significant improvements in numerical performance and policy generalization, providing specific guidance for the rational utilization of expert demonstrations.
Yeor Hafouta, Brendan Williams
Tue 29 Sep
math.PRmath.STstat.TH
pdf html
- reading large, hilbert, expansions, theory
- interest Bayesian computation and structural estimation 0.85
- favorites near “Introduction to regularity structures” 0.85
Abstract
Locally stationary Markov chains have attracted a lot of attention in statistics in the past three decades. In this paper we prove a variety of limit theorems for partial sums generated by such chains. We first provide explicit formulas for the asymptotic mean and variance (and also higher moments), and obtain optimal convergence rates towards them. We then prove a Berry--Esseen theorem, a local central limit theorem, Edgeworth expansions, and large and moderate deviations principles. Our approach involves a parametric Perron--Frobenius theorem, which is proved using the theory of complex (Hilbert) projective metrics developed in \cite{Rugh,Dubois}, together with local approximation arguments and ideas in \cite{dolgopyat2023berry}.
Anna Kazachkova, Leonhard Hennicke, Rainer Schlosser, Ralf Herbrich
Tue 29 Sep
cs.LGcs.AIstat.ML
pdf html
- reading equations, differential equations, pdes, hamiltonian monte
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.84
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.81
Abstract
Prior information for real-world physical quantities is most elegantly expressed via partial differential equations (PDEs). In this paper, we propose a novel way to solve PDEs using probabilistic inference on a factor graph. In general, factor graphs provide a natural way to encode prior knowledge into a model as explicit factors; here, this knowledge is provided by a governing PDE, which narrows the solution space, while observed data further shape the posterior over the parameters. The approximate parameter posterior is inferred using message passing based on moment matching, without posterior sampling or global gradient-based optimization. We demonstrate our approach on the first-order advection and the second-order semi-linear Fisher-KPP equations, where it achieves predictive accuracy comparable to a standard baseline while providing structured predictive uncertainty. Moreover, the inferred posterior marginal means and uncertainty structure match more closely those obtained using Hamiltonian Monte Carlo than the evaluated mean-field variational inference baseline, while requiring up to 10x less training time in our experiments, with inference speed comparable to variational inference.
Md Rezwanul Islam, Wael Mohammed
Thu 24 Sep
cs.LGstat.APstat.ME
pdf html
- reading market, forecasting, gets, business
- favorites near “Measuring Price Selection in Microdata: It's Not There” 0.84
- interest Adaptive data collection and anytime-valid inference 0.85
Abstract
A forecasting benchmark reports which method won. We show that the answer is set by the evaluator's choices before any model is fitted. We benchmark 24 forecasting methods and one textbook reference, including six 2025-era time series foundation models, on a production marketplace panel of 1,887 business customers over 67 months. We hold the data, the horizon and the period fixed, and vary only the evaluation design. Three choices each reverse or dissolve a headline conclusion. Changing the unit of analysis from the market total to the individual customer moves our production baseline from second of nineteen, beaten by nothing, to twenty-third of twenty-five. Nineteen of its twenty-four challengers beat it there. Changing how much error is pooled decides whether a Diebold-Mariano test finds anything at all. Scoring prediction intervals rather than point forecasts reorders the field almost completely, with a rank correlation of 0.02 on intermittent demand. We then measure what the deployed system gets from this. Its selection rule captures 55% of the distance between doing nothing and choosing with hindsight. The reversal is not a quirk of our data. We ran the released protocol, unchanged, on the public M5 retail panel. The same baseline shape places first at the market total and last per series, beaten by everything, and a replayed selection rule closes 64.7% of the same floor-to-ceiling distance there. Adding five zero-shot foundation models to that roster changes who wins at the total, not the shape. The bands' blind spot travels too: conformal bands under-cover most on the spikiest items. Splitting our own panel into ever smaller groups turns the contrast into a curve: the baseline's rank worsens at every level of disaggregation. We release the evaluation protocol and report an error of our own that inverted a result before we caught it.
Michael Grantham
Fri 25 Sep
stat.MEstat.CO
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.83
- interest Adaptive data collection and anytime-valid inference 0.85
- reading special, procedures, functional, imply
Abstract
Most changepoint procedures fix segment forms and boundary relationships before estimating changepoint locations. We introduce GRACE, the General Regime-Aware Changepoint Estimator, for an $\ell_0$-penalized problem that jointly selects the number and locations of changepoints, segment forms, transition types, and continuous parameters. A transition depends on the preceding fit only through the quantities it inherits. GRACE retains accumulated cost conditional on these quantities, using joint envelopes when several must be inherited together. This yields an exact conditional-cost calculus for transitions that preserve, release, or fix finitely many boundary parameters. Optimal partitioning, FPOP, and CPOP arise as special cases. CPOP motivates the construction: its level-conditioned cost supports continuous piecewise-linear fitting but cannot represent a discontinuous level shift preserving the incoming slope, because optimality over endpoint level need not imply optimality over slope. GRACE profiles each generated linear candidate at fixed level and at fixed slope before pruning either envelope. The two envelopes support continuous slope changes, level shifts preserving slope, constant regimes, trend termination and resumption, and complete resets within one exact optimization. Under least squares, finite-dimensional linear-basis models with affine boundary relationships yield quadratic candidate costs. We establish exactness of the general recursion and develop functional, penalty-based, and objective-bound pruning rules. Polynomial and seasonal specializations illustrate the framework, whose recursion permits substantial parallel computation.
Subir Hait
Wed 30 Sep
stat.ME
pdf html
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.89
- interest Adaptive data collection and anytime-valid inference 0.89
- favorites near “Orthogonal Statistical Learning” 0.86
Abstract
Sequential studies with heterogeneous strata often target a weighted population mean while collecting data under a different, possibly adaptive allocation. This creates two distinct design choices: how observations are allocated and whether inference is performed directly for the target or by aggregating simultaneous stratum-level confidence sequences. We compare these architectures under anytime-valid inference. For a single prespecified target, direct target sampling yields one bounded pooled process. When simultaneous stratum-level reporting, post-hoc reweighting, or robustness over several targets is required, a stratified construction aggregates local confidence sequences. For local half-widths with power-law rate n raised to minus beta, we derive the asymptotically width-optimal allocation; its exponent is 1/(1+beta), and root-n variance-adaptive boundaries yield the 2/3 rule. We establish validity under predictable adaptive sampling, show oracle tracking under a vanishing exploration floor, extend the design to uncertain target distributions and familywise best-system identification, and quantify the first-order cost of unnecessary local multiplicity. Simulations and a public benchmark replay show two robust patterns: variance adaptation can matter more than fine allocation tuning, and pooled target-specific inference can reduce stopping cost dramatically when local simultaneous guarantees are not needed. The framework connects stratified sampling, confidence sequences, adaptive allocation, and ranking and selection through a common target-aware sequential design problem.
Ian Arawjo
Wed 30 Sep
cs.CLcs.AIcs.HCstat.ME
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.83
- interest Semiparametric inference and debiased machine learning 0.86
- reading researchers, cis, recommendations, inference
Abstract
Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. First, we find that running statistics over raw LLM judge scores leads to inflated false positives: counterintuitively, for many inter-rater agreement metrics, false positive risk peaks at "almost perfect" human-LLM agreement. To help researchers understand how to run statistics over LLM judges responsibly, we present guidance and tooling for the statistical analysis of mixed human-AI judge designs, and implement nine hypothesis tests via prediction-powered inference (PPI), including the first known PPI corrections for four rank-based tests (Wilcoxon signed-rank, Mann-Whitney U, and omnibus variants). To keep PPI++ stable with small human-labeled calibration sets, we introduce bootstrap-adaptive power tuning, which shrinks the estimated weight toward a target estimated from the labeled data, and accounts for that weight's own sampling variance. Second, through Monte Carlo simulations, we derive recommendations for what CI, p-value, and FWER correction methods to use for small-sample AI evaluations (N<100), and warn researchers against bootstrap CIs. We package these recommendations into evalstats, an open-source Python package that selects calibrated methods automatically, and demonstrate it in three scenarios, including one where a real LLM judge validated at "substantial agreement" would have led a researcher to publish a spurious finding. evalstats is publicly available at https://github.com/ianarawjo/evalstats.
Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Jincheng Ying
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.83
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.87
- interest Reinforcement learning, policy learning and off-policy evaluation 0.83
Abstract
Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schr\"odinger--F\"ollmer Actor--Critic (SFAC), an offline-to-online reinforcement learning (RL) method for Kullback--Leibler (KL)-regularized policy improvement through conditional diffusion. A minimax Bellman critic estimates the advantage function, which defines an exponentially tilted target policy. A Doob $h$-transform expresses this update as a correction to the reference diffusion drift. We derive a posterior-mean representation of the correction and estimate it using paired self-normalized importance sampling (SNIS). Supervised regression on these drift targets updates the neural actor without critic action gradients. In the small-update regime, the KL-regularized update follows the natural policy-gradient direction, and the Doob correction represents the same local change in the space of diffusion drifts. Under suitable conditions, we derive finite-sample bounds that separate the effects of critic estimation, neural drift regression, finite-sample SNIS, diffusion discretization, and inherited actor error on expected average policy suboptimality. Synthetic experiments assess the accuracy of approximation to prescribed advantage-tilted targets and sensitivity to sampling budgets. On six offline-to-online continuous-control tasks, a reference-anchored implementation achieves higher final-window returns than those of its corresponding offline initialization.
Xinyue Xu, Mengtao Wen, Long Feng
Wed 30 Sep
stat.ME
pdf html
- interest Semiparametric inference and debiased machine learning 0.88
- reading eigenvalues, theory, gains, good
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.78
Abstract
Changes in principal eigenvalues and eigendirections provide complementary diagnostics of structural instability in factor models, but covariance-based inference can be unreliable under heavy tails. We develop two-sample tests for equality of these features of trace-normalized shape matrices under high-dimensional elliptical factor models. Our approach combines Tyler's identity with orthogonalization and cross-fitting to accommodate general location and precision pilots. We derive a spectral limit theory that yields asymptotically valid calibration without upper-tail moment assumptions on the radial variables, allowing the two populations to have different radial distributions and factor ranks. Simulations demonstrate good size control and gains in size-adjusted power over covariance-based benchmarks under heavy-tailed elliptical distributions. An application to S\&P~500 stock returns illustrates how the tests distinguish changes in relative component strength from changes in component orientation.
Han Xiao, Yuefeng Han, Rong Chen, Ama Ampadu-Kissi
Fri 25 Sep
stat.MEecon.EMmath.STstat.TH
pdf html
- reading economics, series, large, matrices
- favorites near “Priors for the Long Run” 0.86
- interest Macroeconometrics and time series 0.86
Abstract
Matrix time series is a series of matrix data observed over time. Analytical tools for such time series is needed in many applications in finance, economics, engineering and many other fields. To avoid the use of vectorization of the matrices which loses the column and row information, and the vector autoregression framework in traditional time series analysis, \cite{chen2021autoregressive} proposed the Matrix Autoregressive (MAR) Model. The model maintains and utilizes the matrix structure, leading to a substantial dimensional reduction and admitting explicit interpretations, comparing with the vector autoregressive model on the vectorized data. However, the MAR model still encounters difficulties in dealing with large dimensional matrix time series as the coefficient matrices in MAR models are also large. In this paper we propose to achieve further dimension reduction through reduced-rank constraints of the coefficient matrices in the MAR model. Estimation and rank determination procedures are studied. Theoretical investigation and empirical examples show that the reduced-rank constraint can achieve higher statistical efficiency than the MAR model.
Razieh Nabi, David Benkeser
Thu 24 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.90
- reading supplementary, implications, written, make
- interest Causal inference and identification in econometrics 0.84
Abstract
In causal mediation analysis, natural direct and indirect effects are defined through nested counterfactuals that combine the outcome under one exposure level with the mediator value under another and are therefore inherently cross-world. Their canonical identification additionally relies on cross-world independence assumptions. Consequently, both the estimands and their identifying assumptions remain controversial. In this paper, we ask what additional single-world structure would be required to re-express cross-world estimands as single-world quantities, identifiable under single-world assumptions. We show that these quantities can be written as mixtures of controlled single-world effects under assumptions involving a susceptibility marker, a possibly latent baseline variable that encodes the "would-be" mediator value under a reference treatment arm. If observed, the marker would make several implications of the marker restrictions empirically testable. These results provide a transparent single-world formulation of cross-world mediation effects while making explicit the assumptions required. Although these conditions clarify how cross-world estimands can be interpreted within a single-world framework, their practical relevance depends on whether such susceptibility markers can be justified in specific applications. We discuss implications for principal stratification and show in the Supplementary Material how the formulation extends to ordered mediators and settings with exposure-induced mediator-outcome confounding.
Shinsaku Sakaue
Thu 24 Sep
stat.MLcs.DScs.LG
pdf html
- favorites near “Preference-based Online Learning with Dueling Bandits: A Survey” 0.88
- reading rational, online, feature spaces, feedback
- interest Reinforcement learning, policy learning and off-policy evaluation 0.83
Abstract
We study online inverse linear optimization with a fixed unknown linear utility: in each round, an environment presents a compact action set, the learner recommends an action from it, and the environment returns an action that maximizes the utility over the same set. When the utility vector and the actions lie in the $d$-dimensional Euclidean unit ball, we give a randomized algorithm whose regret---the cumulative utility shortfall relative to optimal actions---is $O(\sqrt d)$ in expectation for every time horizon, without knowledge of the horizon. The dependence on $d$ is optimal up to a constant factor by the known $\Omega(\sqrt d)$ lower bound for horizons $T\ge d$. Our algorithm maintains matrix multiplicative weights on polynomial feature spaces at geometrically spaced scales. It selects a recommendation distribution by solving a linear program and updates its score matrices by comparing the available actions with the feedback action. With rational oracle outputs and feedback actions, an implementation computable relative to a linear-optimization oracle preserves the $O(\sqrt d)$ regret bound. Whether the same rate is attainable with running time polynomial in the dimension, horizon, and input length remains open.
Joris Mulder
Tue 29 Sep
stat.ME
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.82
- reading empirical, comment, literature, resulted
- favorites near “Optimal Decision Rules for Weak GMM” 0.87
Abstract
Vogels et al. (2024) presented an empirical comparison of Bayesian methods for structure learning in undirected Gaussian graphical models. The method implemented in the R package BGGM was run with equal prior probabilities for a null, negative, and positive partial correlation, implying a prior inclusion probability (equivalent to a prior graph density) of 2/3. The other Bayesian methods in the comparison used a prior graph density of 0.2 however. The data-generating densities ranged from 0.01 to 0.1. This resulted in a substantial overestimation of the inclusion probabilities of absent edges by BGGM. For a fair comparison of the performance of the different methods, I repeated the simulation using a prior inclusion probability of 0.2 using BGGM. In this case, the performance of BGGM is comparable with the other Bayesian methods.
Sudarshan Manikantan, Abhishek Bhattacharjee
Mon 28 Sep
stat.MLcs.LGmath.PRmath.STstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.88
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.85
- favorites near “Optimal Decision Rules for Weak GMM” 0.89
Abstract
Adaptive arm selection changes the distribution of the observations collected by a bandit algorithm, but it need not change their limiting empirical spectrum. We study Gaussian bandit designs in which the dimension and the number of observations grow proportionally. A quantitative coupling theorem compares the design generated by any causal selection rule with an independent Gaussian design. If the logarithm of the number of available arms is sublinear in the dimension, the empirical spectral distribution converges to the Marchenko-Pastur law, uniformly over the selection rule. Consequently, Gaussian Bayesian bandits have policy-independent first-order limits for posterior mean-square uncertainty, squared posterior covariance, and information acquisition. For linear-score selection, we obtain the exact conditional arm distribution and show that two-arm selection produces an exactly Wishart Gram matrix in every dimension, despite its nonzero conditional mean. For a fixed selection direction, we identify an explicit eigenvalue and eigenvector transition governed by the second moment of a Gaussian maximum. A counterexample shows that a direction's overlap with a reference signal does not determine this transition. Finally, an exponentially large arm pool permits a different bulk limit, establishing the order-sharpness of the arm-growth condition. These results distinguish global spectral stability from directional effects in adaptive bandit data.
Mehrzad Khosravi, Hema Yoganarasimhan
Mon 28 Sep
econ.EM
pdf html
- reading causal, questions, researchers, inference
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.81
- interest Causal inference and identification in econometrics 0.83
Abstract
Researchers increasingly rely on third-party platforms such as Similarweb and Semrush to measure web traffic when first-party analytics are unavailable. Yet these platforms report model-generated estimates rather than raw data, raising questions about whether their measures preserve the temporal and cross-source variation required for causal inference. As a motivating diagnostic, we examine reported referral traffic around two documented search-engine outages; the absence of visible discontinuities illustrates why preservation of identifying variation cannot be taken for granted. We then characterize three mechanisms, within-source smoothing, cross-source leakage, and treatment-induced calibration error, through which platform processing can generate nonclassical outcome measurement error. Analytical results and a stylized difference-in-differences simulation show that this error can attenuate, amplify, or reverse estimated treatment effects. Our findings caution against using third-party traffic measures based on black-box proprietary models for causal inference.
Patrick Forré
Thu 24 Sep
math.STcs.ITcs.LGmath.ITmath.PRstat.MLstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.88
- reading wealth, inequality, versus, opposite
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.82
Abstract
In safe hypothesis testing with test supermartingales, Ville's inequality provides anytime-valid type-I error guarantees for every significance level $\alpha\in(0,1]$, if one rejects the null hypothesis whenever the wealth process first exceeds $1/\alpha$. Due to an inherent asymmetry, the type-II error behaves differently. We prove two things about the latter, for a simple null and alternative. First, the mean growth rate $\mathbb{E}_{P_1}[\log E]$, the e-power, that Kelly betting and growth-rate-optimal e-variables maximise, bounds nothing on its own. For every level $c>0$, every $\alpha$ and horizon $t$ we construct e-variables of conditional e-power exactly $c$ whose probability of not rejecting by $t$ is arbitrarily close to one. It forces eventual rejection, but no finite-horizon guarantee follows. Second, the quantity that does control the type-II error is the Chernoff-Stein exponent of an e-variable, $\Lambda(E)=\sup_{s\ge0}\{-\log \mathbb{E}_{P_1}[E^{-s}]\}$, whose range is exactly determined: $\sup_E \Lambda(E)=\mathrm{KL}(P_0\|P_1)$, the classical Chernoff-Stein exponent, and so the ceiling of its own per-e-variable form. One conditional application of Hoelder's inequality per step gives it, for every test supermartingale on an arbitrary filtered space, with no independence or product structure; the i.i.d. case adds that it is matched, and attained by nothing. The e-power has its own ceiling, $\mathrm{KL}(P_1\|P_0)$, and that one is attained, $P_0$-a.s. uniquely, by the likelihood ratio $R$. The two optima are the same divergence in opposite arguments, at opposite ends of the flattened family $R^{\beta}/\mathbb{E}_{P_0}[R^{\beta}]$: the ceiling as $\beta\downarrow0$, $R$ at $\beta=1$. Which $\beta$ is best is settled by the horizon, exactly: $R$ is optimal at $t=\log(1/\alpha)/\mathrm{KL}(P_1\|P_0)$ alone, beaten by sharpening $(\beta>1)$ below it and by flattening above.
Haoran Li
Thu 24 Sep
stat.ME
pdf html
- interest Semiparametric inference and debiased machine learning 0.89
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.78
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.84
Abstract
We study structural-change testing in multivariate linear regression when the response dimension is proportional to the sample size and the number of predictors is fixed. Despite its relevance to applications across a broad range of fields, this problem remains underexplored. The alternatives of interest allow multiple unknown changes in a prescribed linear contrast of the predictor effects. We construct a least-squares-based Wald statistic standardized by the residual covariance estimator and scan it over candidate change-point segmentations. We propose flexible segment scans for both single and multiple change points, together with a discretized multiscale scan that reduces the computational cost. When the response dimension and sample size diverge proportionally, we establish weak convergence of the normalized statistic process to a centered Gaussian process, yielding implementable critical values for all three scans. We further characterize their asymptotic power under local alternatives, explicitly describing how the signal, design, and high-dimensional aspect ratio determine the limiting power. The finite-sample performance is examined through simulation studies. We apply the proposed methods to filtered and standardized returns from U.S. equity portfolios to investigate structural changes in their exposures to the Fama--French factors.
Zhifeng Chen, Chenyang Jiang, Yazhen Wang
Mon 28 Sep
stat.MLcs.LG
pdf html
- reading theory, optimization sampling, odes, connection
- interest Computational methods for heterogeneous agent and dynamic economic models 0.82
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
Abstract
Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse diffusion, absent in the reverse process based on ODEs. Second, we incorporate discretization and establish averaged first-order stationarity bounds---the sampling analog of averaged gradient-norm guarantees in nonconvex optimization---for samplers of both overdamped and underdamped diffusion models. As in nonconvex optimization, the convexity-free certificate is local: it guarantees score consistency, not global mode weights.
Stéphane Bouka, Emmanuel De Dieu Nkou, Alban Mbina Mbina, Guy Martial Nkiet
Fri 25 Sep
stat.ME
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.78
- interest Causal inference and identification in econometrics 0.86
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.80
Abstract
This paper introduces a novel semiparametric testing framework for localized point-impact effects within a spatial semi-functional linear regression model. We address an intricate inferential problem that simultaneously accommodates three distinct sources of complexity: an infinite-dimensional global functional slope, an unknown smooth spatial nuisance surface, and data-dependent, unobserved impact locations. To isolate the point-impact coefficients, we construct a robust residualized cross-moment estimator after regularizing the functional component via Tikhonov inverse and estimating the spatial nuisance surface through local-linear smoothing. We establish a multivariate central limit theorem under the null hypothesis of no point-impact effect. We develop two operational test statistics: a general formulation whitened by the long-run covariance matrix to handle arbitrary spatial dependence, and a simpler chi-square calibration valid under short-range spatial dependencies in regression errors. Rigorous finite-sample simulations and an application to Canadian weather data demonstrate the superior power and size control of our methodology.
Prakhar Singhvi, Yi Zou, Abhishek Bhattacharjee
Fri 25 Sep
stat.MLcs.LGmath.STstat.TH
pdf html
- favorites near “Eluder Dimension and the Sample Complexity of Optimistic Exploration” 0.91
- interest Adaptive data collection and anytime-valid inference 0.88
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.84
Abstract
We study Bayesian linear bandits with an isotropic Gaussian parameter, independent Gaussian candidate arms, and Gaussian reward noise when the horizon is proportional to the dimension. The normalized posterior uncertainty has an explicit limit that is uniform over all causal policies. Gaussian posterior identities then determine the limiting parameter overlaps without an assumed closure of the adaptive recursion. These results yield exact regret curves for Thompson sampling, posterior-mean greedy selection, and a family of policies that scale the posterior sampling covariance. The normalized realized cumulative regret converges in L1, uniformly on compact proportional-time intervals. A policy-uniform lower bound identifies the limiting optimal Bayes regret and proves that posterior-mean greedy selection attains it. Thompson sampling incurs a strictly larger leading regret; its instantaneous regret ratio relative to greedy selection lies between one and two and approaches two at long proportional horizons. Closed-form cumulative curves also identify a different comparison in the vanishing-noise limit. Finally, the instantaneous regret converges to a nondegenerate Gaussian decision-loss distribution, rather than to its mean. The analysis separates the amount of information acquired by a bandit policy from the quality of the decisions made using that information.
Jarrod Burgh, Emerson Melo
Mon 28 Sep
econ.EM
pdf html
- interest Reinforcement learning, policy learning and off-policy evaluation 0.86
- reading welfare, policy, aversion, treatment
- your work near “Timing as an Action: Learning When to Observe and Act” 0.80
Abstract
This paper studies risk-averse treatment allocation when individuals self-select into treatment based on unobserved characteristics. We develop a framework that combines the marginal treatment effect approach to endogenous selection with a general class of coherent risk measures that capture distributional preferences over welfare outcomes. We show that the planner's problem admits equivalent interpretations in terms of uncertainty aversion, distributional robustness, and worst-case welfare. For law-invariant coherent risk measures, we derive a Kusuoka representation that expresses the planner's objective as a weighted evaluation of different regions of the welfare distribution and characterize the resulting optimal allocation rule. We further establish finite-sample regret guarantees for empirical risk-averse policy learning, showing how the statistical difficulty of learning a policy depends on the planner's sensitivity to adverse welfare outcomes. The framework nests the risk-neutral policy learning model of \cite{Kiatagawa_Tetenov_2018} as a special case. An application to the \cite{Card1995} college proximity data demonstrates that incorporating risk aversion can lead to economically meaningful changes in optimal college admission policies.
Cici Chen Gu, Eszter Lakatos, Sara Hamis
Wed 30 Sep
stat.MEstat.CO
pdf html
- reading mcmc, programming, probabilistic programming, carlo mcmc
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.81
- interest Bayesian computation and structural estimation 0.84
Abstract
Multimodality is common in many scientific and engineering applications. However, standard Markov chain Monte Carlo (MCMC) convergence diagnostics often penalise multimodality in practice by conflating overall poor mixing with groups of chains that have mixed well within distinct modes. To diagnose multimodality rather than flag it as a sampling failure, we introduce the MCMC convergence graph: a graph-based diagnostic that computes pairwise R-hat values between chains and summarises them in a graph, revealing sets of chains that explore the same posterior mode. We implement the method in the probabilistic programming framework Stan and demonstrate its use on worked examples ranging from regression models with inherent multimodality in the data, to a pharmacokinetic model for which multimodality reflects non-identifiability. The graph-based diagnostic, which we make freely available in the R package mcmcConvergenceGraph and which accepts MCMC output from any sampler, provides both graphical and numerical summaries of MCMC behaviour for diagnosing multimodality.
Md Rezwanul Islam, Wael Mohammed
Thu 24 Sep
cs.LGecon.EMstat.AP
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.80
- interest Macroeconometrics and time series 0.86
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.81
Abstract
Global forecasting models pool many series and learn one shared function. Gradient-boosted trees are their most common form. We measure a failure of this design that has not, to our knowledge, been documented. Train a global tree on the individual series of a hierarchy, then ask it for the hierarchical aggregate. The aggregate sits far outside the model's training range, and the forecast collapses. The model under-predicts the total by 30-50x in our production deployment, and by up to 496x in a public M5 reconstruction. The mechanism is known: beyond its training range, a tree predicts a constant. It surfaces at the aggregate because the total dwarfs every training series. The cure is not new. Per-series scaling, the preprocessing step that Montero-Manso and Hyndman (2021) recommend, prevents the collapse. So do a weighted aggregate-level training row and seasonal differencing. Our contribution is the characterization. The collapse reproduces on five panels: a production business-to-business marketplace, a synthetic hierarchy, M5, Australian Tourism, and a public business-buyer panel. It holds on three tree libraries, is invariant across training seeds, and is statistically significant. Its onset is immediate and tracks a simple support bound: a scale gap of only 1.15x already costs a third of the total. No standard configuration change prevents it: pooling every hierarchy level into training fails at scale, and the one knob that fits linear models in the leaves softens it without curing it. Rolling the forecasts forward recursively separates the cures: the aggregate-row cure re-collapses, per-series scaling degrades but stays low, and only seasonal differencing keeps its one-step accuracy unchanged. We close with a three-step procedure for diagnosing and preventing the failure in deployed systems.
Isfandiyor Akhmedov
Fri 25 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.88
- interest Causal inference and identification in econometrics 0.84
- favorites near “Difference-in-Differences with multiple time periods” 0.82
Abstract
Covariates summarized over a subject's completed follow-up are sometimes entered into Cox regression as though observed at baseline. This practice incorporates future event or censoring information and changes both the estimand and its sampling behavior. We analyze an affine class in which a genuine baseline covariate is contaminated by realized follow-up time. Treating partial likelihood as an observed-data estimation criterion, we derive the population score and characterize its unique least-false coefficient. Exact continuous-time benchmarks show scale reduction and saturation under strong contamination, while administrative censoring destroys the reduction and may produce overshoot. Grouping exit times with Breslow ties changes the geometry: the coefficient has a single hump and eventually returns to zero even though the induced association diverges. We also derive an observed-data influence function and show why model-based variance can be either too small or too large. A subject-level sandwich consistently estimates uncertainty around the least-false target under the stated conditions, but it does not correct the target itself.
Stéphane Galatolo, Stéphane Chrétien
Mon 28 Sep
stat.MLcs.LGmath.DS
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.84
- favorites near “Recent Developments in Machine Learning Methods for Stochastic Control and Games” 0.82
- reading include, stochastic gradient, large, algorithms particular
Abstract
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most rudimentary, while still outperforming more recent algorithms at a number of learning tasks. A major open question about the current methods used in deep learning is to understand their convergence properties. Following a line of previous works about the long time behavior of gradient-type algorithms, %and in particular the recent contributions from Azizian et al., we present a new approach for studying the asymptotic properties of a wide family of methods from an ergodic theoretical viewpoint. Our main results include a study of the expected time for a stochastic optimisation algorithm to reach a certain small neighborhood of a minimizer and show that this reaching time distributes exponentially around its average, which is given by the inverse of the stationary measure of the target. The assumptions on the Stochastic Gradient noise include the Gaussian and the Sub-Exponential assumptions.
Gon Buzaglo, Elad Hazan
Tue 29 Sep
cs.LGstat.ML
pdf html
- reading infinite, online, regret, adaptively chosen
- interest Online learning and statistical learning theory 0.88
- favorites near “The Minimax Regret of Sequential Probability Assignment, Contextual Shtarkov Sums, and Contextual Normalized Maximum Likelihood” 0.87
Abstract
We consider contextual binary prediction with i.i.d. contexts from an unknown distribution and adaptively chosen losses. We show that a simple Follow-the-Perturbed-Leader algorithm with Gaussian perturbation for each observed context achieves the optimal $\widetilde O(\sqrt{T\log N})$ expected regret for a class of $N$ experts, while requiring one optimization-oracle call per round and no explicit enumeration of the class. For an infinite hypothesis class $\mathcal H$, the algorithm attains $\widetilde O(\sqrt{T\operatorname{VC}(\mathcal H)})$ regret. This resolves an open problem posed by Lazaric and Munos (2012), showing that hybrid classification is computationally as easy as statistical learning.
Tomoshige Nakamura, Ryo Emoto
Wed 30 Sep
stat.MEmath.STstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.89
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.88
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.85
Abstract
In factorial randomized experiments with $K$ binary treatment components, we study which baseline covariates modify component main effects and interactions. We define effect-modifier cells $(j,S)$, $j\in[p]$ and $\emptyset\neq S\subseteq[K]$, through the coefficients $\theta^\ast_{jS}$ of population linear projections of conditional Walsh--Hadamard contrasts $\tau_{S}$ onto standardized covariates. Using baseline-robust direct scores, we develop frequentist inference for all $p(2^K-1)$ cells. First, we prove a debiased central limit theorem with global insensitivity: the residualization baseline enters the cross-fitted score through an exactly conditionally mean-zero perturbation. The influence-function remainder therefore requires only foldwise weighted $L_2$ consistency of the baseline estimator, with no prescribed polynomial rate or nuisance product-rate condition. A separate, explicit penalty-scale condition controls the baseline's effect on the score Lasso. Second, under explicitly stated uniform nuisance, baseline, studentization, and influence-array conditions derived from bounded sparse-regression primitives, we establish a high-dimensional Gaussian approximation over the full grid, calibrated by a Walsh-spectral multiplier bootstrap. This yields simultaneous confidence bands and Romano--Wolf step-down selection of modified cells with strong familywise error control, without effect-heredity assumptions. The factorial design is fixed in the asymptotic analysis, while the covariate dimension may grow; all $2^K-1$ contrasts are covered jointly.
Subir Hait
Fri 25 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.86
- interest Causal inference and identification in econometrics 0.84
- favorites near “Cause, Effect, and the Structure of the Social World” 0.82
Abstract
Researchers often summarize mediation in clustered data with a single-level product-of-coefficients estimator. We show what this pooled analysis estimates and why the target depends on the design. The pooled a-path is weighted by the within-between variance composition of the treatment. The pooled b-path uses a different weight, based on the mediator variance that remains after residualizing on treatment and covariates. Because these weights generally differ, the pooled indirect effect can differ from a level-respecting combination of the within- and between-level indirect effects. We derive the pooled probability limit and split this difference into a nonlinearity term and a weight-mismatch term. Equal b-paths remove the distortion, but equal a-paths do not. Allowing the two weights to vary gives a bilinear surface with saddle geometry, sharp corner bounds, and a zero-distortion frontier. We also define a design-standardized indirect effect and practical diagnostics. Simulations verify the identities and show that a common decompose-first, cluster-weighted construction can overstate apparent distortion. In an ECLS-K:2011 illustration with 8,363 children in 753 baseline schools, the within-level weights were 0.723 and 0.896, a difference of -0.173 (95% bootstrap CI [-0.193, -0.151]). The conventional construction produced 1.76 times the absolute distortion of the identity-consistent construction. The identity-consistent AD interval included zero. Thus, the application shows unequal weighting and sensitivity to construction, not evidence of nonzero population distortion.
Zihao Chen
Mon 28 Sep
stat.MEecon.EMstat.AP
pdf html
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.89
- interest Adaptive data collection and anytime-valid inference 0.87
- favorites near “Dynamically Optimal Treatment Allocation” 0.82
Abstract
Online experiments must often be evaluated before long-term outcomes mature. Under rolling enrollment, these outcomes are observed only for early enrollees, while short-term surrogates are available for everyone. We compare seven estimators across eleven data-generating processes, spanning partial mediation, drift, outcome sparsity, and enrollment-time labeling, with up to $R = 2,000$ replications over more than 500 method-by-scenario cells. We find a sharp robustness-efficiency tradeoff: the surrogate index delivers large efficiency gains when surrogacy holds but its coverage collapses under violations, while PPI-family methods stay asymptotically valid under random labeling at smaller gains. We give the finite-sample variance of PPI++ in the all-units parameterization for a fixed predictor, a joint asymptotic distribution for the two estimators under cross-fitting, and a Hausman-type estimator-disagreement diagnostic, then quantify the detection-damage gap: in the partial-mediation design, where we locate both edges, a band of violations destroys surrogate-index coverage yet is too small to detect on most datasets. On the 64,000-customer Hillstrom experiment the diagnostic rarely flags a violation that biases the surrogate index, and a Cauchy-kernel hybrid of the two estimators inherits 10.7% relative bias; on the 14-million-user Criteo experiment the violation is detected, and subsampling traces detection turning on with scale as damage persists. PPI++ has limits: at a rare-conversion $n = 30,000$ Criteo subsample its empirical coverage is 85.0%. We recommend prespecifying PPI++ with the exact variance as the primary analysis under random labeling and adequate labeled outcome counts, reading the diagnostic as a warning, not a certificate, and treating the surrogate index as a sensitivity analysis.
Kewen Pan, Ying Tang
Thu 24 Sep
cs.LGcond-mat.dis-nnstat.ML
pdf html
- reading operators, variational, govern, diffusion models
- interest Computational methods for heterogeneous agent and dynamic economic models 0.80
- favorites near “Priors for the Long Run” 0.81
Abstract
Conventional score-based diffusion models learn scores without representing normalized densities, whereas tractable normalized models support both sampling and direct likelihood evaluation. A recent tensor-network approach provides such a representation but is largely restricted to low-dimensional lattices. Here we introduce a discrete diffusion model that parameterizes normalized probability distributions using variational autoregressive networks. Explicit Markov jump operators govern the forward noising and reverse denoising dynamics, extending discrete diffusion models with normalized distributions to spin systems on higher-dimensional lattices. We apply this framework to the two- and three-dimensional Ising models across ordered, critical, and disordered regimes, accurately computing thermodynamic quantities including free energy, energy, and magnetization. We further integrate the framework with Monte Carlo sampling, using adaptive diffusion steps to maintain high acceptance rates even at low temperatures while enhancing sample diversity. These results establish a neural-network framework for the discrete diffusion model with normalized probability distributions.
Calvin Tolbert
Fri 25 Sep
stat.MLcs.LG
pdf html
- interest Semiparametric inference and debiased machine learning 0.88
- your work near “Valid Inference with Imperfect Synthetic Data” 0.83
- favorites near “Variational Bayes under Model Misspecification” 0.87
Abstract
Matrix completion, the problem of estimating missing entries in a matrix from noisily observed ones, underlies a diverse array of problems such as recommender systems and counterfactual outcome estimation in panel data. Many algorithms address the problem using regularized least squares, often with the nuclear norm as a regularizer, but this method yields a point estimate with no built-in uncertainty quantification. A Bayesian formulation is a natural alternative, and if the noise variance is known, the nuclear norm-based prior yields a log-concave posterior. Unfortunately, in practice, the noise variance will not be known a priori, so for a fully Bayesian approach, a prior must be imposed on it. We give the first sampler for this model with an explicit non-asymptotic guarantee: polynomial in the matrix dimensions and in the reciprocal of the target accuracy. Our technique is to discretize the distribution of the noise precision onto a grid and build a categorical posterior via thermodynamic integration. This extension is not specific to matrix completion and may be useful in other non-log-concave sampling problems where the non-log-concavity is restricted to a single variable and the joint distribution of the remaining variables is nonsmooth. Our contribution is a feasibility result: we show that a polynomial-time Bayesian sampler for this model exists at all, and the resulting complexity, while polynomial, is not intended as a deployable algorithm at current problem scales.
Andrew McInerney
Fri 25 Sep
stat.ME
pdf html
- interest Semiparametric inference and debiased machine learning 0.89
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.87
- your work near “Valid Inference with Imperfect Synthetic Data” 0.81
Abstract
Variable selection using information criteria has an explicit statistical target but requires discrete search over candidate models. The smooth information criterion (SIC) replaces the discontinuous model-dimension term by a differentiable approximation, with an $\epsilon$-telescoping continuation strategy progressively sharpening this approximation without data-driven selection of a regularisation-strength parameter. We develop SIC as a general procedure for coefficient-level variable selection in generalised linear models (GLMs) and use it to address a central question: how faithfully does smooth optimisation reproduce the corresponding discrete information-criterion selection problem? Focusing on BIC, we benchmark SIC directly against exhaustive subset selection where feasible, using exact support agreement, BIC difference and selection behaviour across a varying signal-strength boundary. Simulations in Gaussian, binomial and Poisson regression show that SIC closely reproduces exhaustive BIC selection and tracks the exact BIC selection boundary. Relative to stepwise BIC, LASSO, SCAD and MCP, SIC produces competitive variable-selection performance while retaining sparse models, with predictive performance broadly comparable across methods. Computational advantages over stepwise BIC increase with predictor dimension. In a real-data application with 16 candidate predictors, SIC recovers the globally BIC-optimal model among all 65,536 supports at a small fraction of the computational cost of exhaustive enumeration.
Junichi Hirukawa, Masanobu Taniguchi, Marc Hallin
Tue 29 Sep
math.STstat.TH
pdf html
- favorites near “Optimal uniform convergence rates and asymptotic normality for series estimators under weak dependence and weak conditions” 0.91
- reading regular, series, class, asymptotic efficiency
- interest Macroeconometrics and time series 0.85
Abstract
For a class of regular estimators, H\'{a}jek, in his celebrated ``Convolution Theorem,'' showed that the asymptotic distribution of a regular estimator is the convolution of the distribution of an efficient estimator and some residual distribution. This result constitutes the foundation of the concept of asymptotic efficiency of regular estimators. In this paper, we provide a second-order version of that classical result. Introducing a class of second-order regular estimators with a valid Edgeworth expansion, we derive their asymptotic distribution under contiguous alternatives and show that it is the convolution of the second-order efficient distribution and some second-order residual distribution. This constitutes a second-order extension of H\'{a}jek's convolution theorem. Based on this, we introduce a concept of {\it second-order robustness} for second-order regular estimators. For a class of general Bayes estimators and minimum contrast estimators in time series models, this second-order robustness is used (i) in the characterization of second-order robust priors, (ii) in a comparison between the second-order robustness of maximum likelihood and Whittle estimators.
Xiang Wang, Kexuan Li, Lingli Yang, Weidong Ma
Tue 29 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.88
- interest Semiparametric inference and debiased machine learning 0.86
- favorites near “Difference-in-Differences with multiple time periods” 0.82
Abstract
Time-to-event outcomes are often collected together with auxiliary outcomes that may provide additional information about baseline risk. We develop a semiparametric approach for incorporating such information into Cox regression without specifying a joint likelihood. The survival outcome follows a Cox proportional hazards model and a continuous auxiliary outcome follows a partially linear model, with separate nonparametric baseline covariate effects linked through a quadratic penalty. When these effects coincide and the outcome scores satisfy the information identities and a first-order orthogonality condition, we derive the sandwich covariance of the penalized estimator and show that, for any fixed penalty level, the asymptotic variance of the Cox regression estimator is no greater than that under separate estimation. We further characterize departures from the shared-effect setting. Local differences of order \(n^{-1/2}\) induce an explicit mean shift in the limiting distribution, yielding a direct bias--variance trade-off, whereas fixed differences generally alter the population target under nonvanishing penalization. These results motivate an adaptive penalty that borrows information when the fitted covariate effects are close and approaches separate estimation when a persistent difference is detected. Simulations and a real world data analysis demonstrate the validity and effectiveness of the proposed method.
Muluneh Alene, Kelly Van Lancker, Stijn Vansteelandt
Wed 30 Sep
stat.MEstat.AP
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.89
- interest Semiparametric inference and debiased machine learning 0.84
- reading adjustment, standard errors, randomized, treatment
Abstract
In contemporary randomized controlled trials (RCTs), the number of patients is often small relative to the number of baseline covariates collected. In such settings, maximum likelihood estimators of conditional treatment effects in generalized linear models, along with their standard errors, may exhibit substantial bias. This study examines whether similar bias arises in G-computation estimators of marginal treatment effects, which are known for their robustness to model misspecification when applied with canonical GLMs. We develop theoretical insights, drawing on recent literature on G-computation under proportional asymptotic regimes, in which the number of covariates grows with the sample size. Specifically, we characterize the bias of the standard G-computation estimator and the leave-one-out G-computation estimator under such high-dimensional settings. Monte Carlo simulations using linear and logistic outcome models are conducted to evaluate practical remedies for G-computation estimators, including covariate selection, cross-fitting, leave-one-out cross-fitting, and small sample corrections to standard errors. Further insights are derived from a re-analysis of the BestAIR trial data. Our findings provide guidance for the application of G-computation in modern RCTs, particularly when addressing challenges posed by limited sample size. \textbf{Keywords:} Covariate adjustment, covariate selection, high-dimensional, proportional asymptotics, targeted learning, cross-fitting.
Bram Wouters, Cees Diks
Thu 24 Sep
stat.MEq-fin.ST
pdf html
- interest Macroeconometrics and time series 0.87
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.80
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
Abstract
We develop a model-agnostic framework for noise reduction in high-dimensional time series that explicitly targets optimal recovery of a low-dimensional latent dynamic component contaminated by observational white noise. Under the assumption that the latent dynamics live in a low-dimensional linear dynamic subspace, we characterize the optimal linear projection onto the dynamic subspace and provide a geometric description of the residual error in terms of the relative orientation of the signal and noise spaces. We propose estimators for the dynamic subspace and the optimal projection based on lagged covariance matrices, bootstrap dimension selection, and a low-rank representation of the structured noise. Under mild conditions, the resulting denoised series is shown to converge to its population target at the usual parametric rate. Simulations show that the proposed method can substantially improve subspace estimation, reconstruction error, and one-step-ahead forecast accuracy compared with both orthogonal projection-based denoising and the raw data. The approach is illustrated by empirical applications to high-dimensional stock returns and to a 20-variate time series of macroeconomic indicators.
Yi-Han Wang, Peng Zhao, Zhi-Hua Zhou
Tue 29 Sep
cs.LGmath.OCstat.ML
pdf html
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.87
- your work near “Timing as an Action: Learning When to Observe and Act” 0.81
- interest Adaptive data collection and anytime-valid inference 0.84
Abstract
Non-stationary online learning has attracted much attention in recent years, as static regret is insufficient to guide algorithm design in changing environments. To address this limitation, interval regret and dynamic regret have been introduced as two representative performance metrics that strengthen static regret in complementary directions. Interval regret requires an online algorithm to achieve competitive static regret over every local time interval, whereas dynamic regret evaluates performance against an arbitrary sequence of time-varying comparators. Despite their importance, the relation between these metrics has long remained unclear. Prior work has often regarded interval regret as the stronger notion, based on the intuition that local guarantees should naturally induce global guarantees. Consequently, it is widely conjectured that an algorithm with optimal interval regret should automatically attain optimal dynamic regret. In this paper, we first establish a negative result that refutes this intuition of a metric-level implication. Specifically, for both convex and curved functions (including exp-concave and strongly convex functions), we show that there exist instances in which an algorithm with optimal interval regret nevertheless fails to achieve optimal dynamic regret. We then show how to leverage local adaptivity to obtain optimal dynamic regret. In particular, optimal dynamic regret can be attained by invoking an interval regret minimization process over an enlarged Euclidean ball containing the original convex feasible domain and using a suitable domain-converted surrogate loss. This reduction applies to both convex and curved functions. As a byproduct, we obtain the first proper and efficient algorithm with optimal dynamic regret for exp-concave functions, improving prior results while significantly simplifying the analysis.
Xuanyu Chen, Yue Yu
Tue 29 Sep
stat.MLcs.LGmath.STstat.TH
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- interest Online learning and statistical learning theory 0.87
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.82
Abstract
We study smoothed online prediction with bounded adversarial responses. This widely studied framework bridges i.i.d. sampling and adversarial covariate selection through a smoothness parameter $\textsf{C}_{\textsf{cov}}$, which bounds conditional covariate densities relative to a fixed, unknown base measure. We propose \textsc{Hedge-Cover}, an information-theoretic algorithm that achieves sublinear regret $\widetilde{O}(\sqrt{\text{Pdim}(\mathcal{F}) \textsf{C}_{\textsf{cov}} T})$ for function classes with bounded pseudo-dimension. The algorithm aggregates a carefully constructed family of experts using \textsc{Hedge}, with a prior that links regret to the number of disagreements between a consistent selector and a target function. We bound this number by exploiting covariate smoothness. This answers an open problem posed in \cite{blanchard2025agnostic} on the minimax optimal adaptive regret of the smoothed online regression problem. We establish a matching lower bound for the class of linear predictors. The main intricacy of the lower bound lies in explicitly constructing a challenging sequential covariate distribution supported on mutually orthogonal hyperplanes. This construction may be of independent technical interest. Finally, we revisit the well-specified setting and quantify the effect of response noise. For conditionally $\nu^2$-subGaussian responses, we extend the existing lower bound under realizable responses by showing that the minimax expected regret is $\Omega((1\vee \nu)\sqrt{(\textsf{C}_{\textsf{cov}}-1)dT})$ for a function class of VC dimension $d$. A corresponding upper bound for ERM matches this dependence on $\nu$.
Matthieu Lerasle, Antoine Roche
Tue 29 Sep
math.STstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.86
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
- reading existence, theory, pac-bayes, empirical
Abstract
We study the maximum likelihood estimation of the coefficient {\beta} in well-specified Poisson regression. Using tools from empirical process theory and random conic geometry, we show that the probability of existence of the maximum likelihood estimator (MLE) exhibits a sharp phase transition at the threshold n > d. We then determine a minimum threshold exponential in the norm of {\beta} on the sample size n to guarantee with high probability an excess risk of the asymptotic order d/n. We reveal the existence of an intermediate regime in Poisson regression, when n is larger than d but smaller than this exponential threshold, where the MLE exists but does not achieve the optimal rate d/n. We close the gap between the two regimes up to a d^{1+\epsilon} term with {\epsilon} \in (0, 1) by providing an upper bound on the distance between the MLE and {\beta} whenever n > d^{1+\epsilon}. Along the way, we provide two generalizations of well-known PAC-Bayes inequalities regarding sub-Gamma random vectors and sub-Gamma random matrices that are of independent interest and that we use extensively to prove the main results of the present paper.
Patrick Forré
Thu 24 Sep
math.STcs.ITcs.LGmath.ITmath.PRstat.MLstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.87
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.83
- reading wealth, moment, inequality, lead
Abstract
In safe hypothesis testing with test supermartingals, Ville's inequality provides anytime-valid type-I error guarantees for every significance level $\alpha\in(0,1]$, if one rejects the null hypothesis whenever the wealth process first exceeds $\frac{1}{\alpha}$. Due to an inherent asymmetry, the type-II error does not have such guarantees: a heavy concentration of the probability on the lower tail of the log-increments can lead to one catastrophic bet that undoes any amount of accumulated evidence. This paper studies how different hypotheses on those lower-tail probabilities lead to different bounds on the type-II error of the sequential test. They all reduce to one master inequality, which bounds the type-II error at level $\alpha$, at a fixed horizon and sequentially, in terms of a one-sided Legendre transform of the (inverse-)moment generating function of the e-variables, evaluated at one number: the amount by which the lower bound of the accumulated e-powers exceeds $\log\frac{1}{\alpha}$. And, the step is lossless, in the sense, that it extracts exactly a constrained information projection. Every bound presented here is a corollary, obtained by a certain majorant of the above function. The hypotheses are: a finite negative moment; an exponentially small crash probability with a moment on the winning side; a wealth floor with a conditional variance, and its Bernstein variant, which interpolates between a Gaussian regime set by the variance and an exponential one set by the scale; a sub-Gaussian or bounded-tilt lower tail; bounded log-increments; and i.i.d. increments, where the majorant is the truth. We also provide an empirical-Bernstein variant. Each hypothesis may either be read as a condition on the e-variables one has, or as the price of betting with an approximation to the likelihood ratio rather than the ratio itself, which satisfies the weakest condition for free.
Karolina Gliszczyńska-Schroeder
Fri 25 Sep
econ.EM
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.89
- interest Semiparametric inference and debiased machine learning 0.84
- reading propensity score, treatment, propensity, units
Abstract
Studying heterogeneous treatment effects has become essential in experimental and observational studies. A critical assumption for obtaining reliable treatment effect estimates is overlap, which requires that treated and control units have sufficiently similar covariate distributions. Poor overlap may limit the effectiveness of estimators, especially those based on propensity scores, potentially leading to unreliable results. We investigate the effectiveness of kernel balancing (KBal) (Hazlett, 2020) as an alternative to propensity score methods for conditional average treatment effect (CATE) estimation, particularly in settings with overlap violations. Building on optimization-based balancing approaches, we integrate KBal weights into tree-based methods, specifically, causal forests (Athey et al., 2019) and the X-Learner (XRF) (K\"unzel et al., 2019), to assess their impact on bias reduction and estimation precision. Monte Carlo evidence shows that KBal achieves near-exact balance in a transformed feature space, thereby improving treatment effect estimation in cases where traditional reweighting methods struggle due to extreme weights, finite-sample bias, or insufficient removal of pre-existing confounding bias. We apply the proposed methods to the semi-synthetic IHDP benchmark dataset. Overall, the results indicate that KBal leads to performance improvements, especially in settings with nonlinear treatment effects and limited overlap, making it a useful alternative to propensity score methods.
Amitakshar Biswas, Yuhan Li, Ruoqing Zhu
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.82
- interest Reinforcement learning, policy learning and off-policy evaluation 0.83
- favorites near “Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic Policies” 0.88
Abstract
Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing weight. In ideal settings, this mixed Bellman formulation can provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting. To solve this mixed residual optimization, we adopt a minimax formulation involving a critic function. Unlike standard approaches that rely on a fixed functional class, we construct a data-dependent critic representation using predicted future feature directions which effectively induces a kernel adapted to the underlying transition dynamics. This allows the critic to focus on directions that are most relevant for the estimated Bellman error. To control overfitting, we use sample splitting to construct the critic and estimate the value function on separate data subsets. Simulation studies and MetaWorld tasks illustrate the effect of the mixing parameter and show that intermediate residual combinations can improve value estimation in challenging settings.
Romane Giard, Yohann De Castro, Roland Denis, Clément Marteau
Fri 25 Sep
math.OCmath.STstat.MLstat.TH
pdf html
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.88
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.83
- interest Kernels, operators and functional data 0.84
Abstract
This paper investigates the numerical resolution of the Beurling-LASSO (BLASSO), a convex optimization framework that promotes sparsity in the space of measures. We consider its application to the estimation of Gaussian mixture models (GMMs) with an unknown number of components and unknown diagonal covariance matrices. Our approach combines the Conic Particle Gradient Descent (CPGD) principle with Riemannian gradient descent, to account for the underlying Fisher-Rao geometry of Gaussian distributions. Our contributions are twofold. First, we provide theoretical guarantees for the convergence of our algorithm. In particular, we establish exponential local convergence under a non-degeneracy condition on the solution and relate this assumption to a separation condition on the underlying statistical target. Second, we address practical implementation aspects of CPGD and present numerical experiments illustrating its performance. On the test cases considered, these experiments suggest that CPGD is more robust to overspecification of the number of components than the EM algorithm. We also investigate the impact of component separation on recovery accuracy.
Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Jiequn Han, Lexing Ying
Wed 30 Sep
cs.LGstat.COstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.85
- interest Computational methods for heterogeneous agent and dynamic economic models 0.82
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
Abstract
Many inference-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the pretrained distribution. Feynman-Kac sequential Monte Carlo (SMC) makes this correction exact in principle, but its prescribed weights routinely degenerate when the proposal dynamics are misaligned with the tilt, capping the practical gains from additional particles. We introduce FluxLite, a lightweight, training-free proposal-control framework for discrete diffusion. On the sparse directed graph of pretrained reverse rates, any sparse jump-rate perturbation can be exactly compensated by a $q_t$-weighted graph-divergence term in the Feynman-Kac potential; the target path is therefore preserved while the residual reweighting variance becomes a local convex objective. We instantiate this principle as two practical samplers: a one-hop local reallocation rule (HEU) and a small nonnegative quadratic program over pretrained-rate bases (D-VCG). We further prove population stability under the standard score-entropy training loss, identifying a tilted-path coverage factor that governs robustness to score error, together with finite-particle convergence for a fixed controlled Feynman-Kac recursion. Empirically, FluxLite improves over standard Feynman-Kac SMC baselines by up to two orders of magnitude in terminal KL on an analytically tractable finite-state CTMC benchmark, and reduces row-correlation MSE on 2D Ising sampling by 5-7x in geometric mean and up to 55x at peak.
Florian Gach, Simon Hochgerner
Mon 28 Sep
math.STmath.PRq-fin.MFstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.86
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.79
- reading control, globally, given, scale
Abstract
This article is concerned with the asymptotic shape of quantile surfaces, defined as the set of quantiles at a given level $\alpha$ generated by a controlled one-dimensional distribution. Specifically, when the distribution arises as a linear combination of log-normal random variables and the control is a vector of positive coefficients, we prove that quantile surfaces are globally concave in the left tail ($\alpha\to0$) and globally convex in the right tail ($\alpha\to1$). Moreover, these surfaces exhibit asymptotic separation of scale and shape.
Alexey Kurennoy
Mon 28 Sep
stat.ME
pdf html
- reading units, theory, connections, treatment
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.82
- interest Reinforcement learning, policy learning and off-policy evaluation 0.81
Abstract
Bipartite A/B tests are experiments in which treatment is randomised over one set of units, while outcomes are measured on another. For example, an online marketplace may test a new pricing algorithm on a random subset of items, while the outcome of interest, say, purchase satisfaction, is measured on customers, each of whom interacts with many items. Existing methods for analysing bipartite experiments assume that every randomisation unit is assigned to treatment or control. In practice, often only a subset participates: platforms cap rollout risk, reserve holdout groups, and split their population across concurrent tests. Ignoring the unassigned units biases estimation, while including them requires care. We construct an unbiased linear estimator for bipartite experiments with partial assignment. Observations must be reweighted not only by the (centred) share of treated connections among participating ones (the exposure) but also by the number of participating connections, each inverse-weighted by its participation propensity, so that units whose connections are well covered carry proportionally more weight. The resulting estimator, EARL (Exposure- and Allocation-Reweighted Linear), is unbiased, consistent, and asymptotically normal; we devise two asymptotic variance estimators and show that it has minimal variance in a natural class of linear estimators. EARL remains unbiased regardless of the experience the unassigned units receive, as long as they contribute to expected outcomes additively; in particular, they can be allocated to other, non-overlapping tests. Our theory is complemented with a simulation study on two public datasets, in which EARL attains up to six times lower error than the strongest existing baseline and, in some configurations, over an order of magnitude lower than Horvitz-Thompson-style alternatives.
Stefan Oberdörster
Thu 24 Sep
math.PRcs.LGstat.COstat.ML
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.82
- interest Bayesian computation and structural estimation 0.85
- reading equilibrium, advances, slower, rates
Abstract
Local geometric mixing localizes geometric mixing by requiring geometric convergence to equilibrium in total variation only over finitely many transitions. It accommodates local convergence rates and captures rapid local equilibration, even when global mixing is much slower. We establish and discuss local geometric mixing bounds through Dobrushin contraction. We then apply this approach to Diffusion Path Monte Carlo, a recently proposed Markov chain Monte Carlo method, aimed at leveraging advances in score-based modeling, whose ideal transitions coincide with those of the Proximal Sampler. Our analysis covers both the ideal method and its implementable Metropolis-adjusted counterpart, providing mixing guarantees under minimal assumptions. For the ideal method, these guarantees complement recent spectral gap estimates, which we develop into mixing time bounds.
Keyu Li, Ruoxu Tan
Thu 24 Sep
stat.MEstat.ML
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.83
- reading operators, theory, functional, causal
- interest Causal inference and identification in econometrics 0.83
Abstract
We study causal discovery where each node is a random function. Previous studies on this topic rely on structural assumptions, e.g., linearity or non-linearity, and distributional assumptions, e.g., Gaussianity or non-Gaussianity. In contrast, we make use of covariance operators to avoid these assumptions. Under functional additive noise models, we propose a new sufficient condition to identify a valid topological ordering based on comparing norms of conditional covariance operators. Taking advantage of this identifiability condition, we develop a new mixed regression model that subsumes linear and non-linear models. Together with variable selection, our procedure yields an estimation of the causal directed acyclic graph (DAG) for functional variables. In theory, we develop the least-squares-type theory of this regression model, and derive asymptotic consistency of order determination, sparse regression, as well as identifying the DAG. Computational algorithms based on discrete observations are provided. Applied to simulated data, our approach performs satisfactorily among existing approaches. A real data example of brain effective connectivity is also presented.
Paweł Lenartowicz, Hubert Plisiecki
Wed 30 Sep
cs.LGstat.ML
pdf html
- reading science, inference, spectrum, basis
- interest Semiparametric inference and debiased machine learning 0.85
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.84
Abstract
Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional X and is widely used across applied science, but inference on the resulting fit is either expensive, biased and discouraged, or absent. We reduce inference to held-out OLS refits of the supervised subspace, a primitive shared by PLS, supervised PCA, and linear probes, and supply two tests using held-out correlations: a Nadeau-Bengio corrected asymptotic t-test as a fast approximation, and a permutation test with comparable power, finite-sample valid under outcome-predictor independence and iid rows. Held-out predictions are unchanged under any orthogonal rebasing of the supervised span, so an interpretable basis such as varimax inherits the joint claim but not a per-axis p-value; per-component claims come from a fixed-sequence test on the PLS extraction order. We validate on synthetic geometries, two NIR chemometric datasets, and cross-lingual word-embedding regressions; the exact test also transfers to supervised PCA and a ridge probe. The proposed tests have more power than CV-permutation-Q^2, at a fraction of its cost. A pre-run check on n and the spectrum of X says when the approximation is safe. We release a Rust library with Python, R, and Julia bindings, plus a Python text pipeline.
Apoorv Srivastava, Eric Darve
Mon 28 Sep
cs.LGstat.ML
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.83
- favorites near “Spectral State Space Models” 0.79
- interest Bayesian computation and structural estimation 0.85
Abstract
Sequential inference estimates latent states from noisy and incomplete observations. Particle Filters (PFs), a class of Monte Carlo methods based on importance sampling, provide a flexible framework for this task, but often suffer from poor sample efficiency and unfavorable scaling with dimension, partly due to suboptimal proposal distributions. We address these challenges by integrating learned proposals into the PF framework. We introduce Neural Optimal Particle Filters (NOPFs), which learn an amortized approximation to the optimal proposal from offline simulated one-step conditioning tuples. The learned proposal is used as a drop-in replacement in standard PF updates, with samples corrected by standard importance weights so that the method asymptotically targets the same filtering distribution under standard support and density-evaluation assumptions. Across stochastic nonlinear benchmarks of varying inference complexity, NOPFs improve sample efficiency and distributional accuracy over standard PF baselines with modest computational overhead. The approach integrates data-driven proposal learning into classical inference without altering the underlying filtering objective.
Matteo Grigolettto, Francesco Lisi
Tue 29 Sep
stat.ME
pdf html
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.83
- interest Macroeconometrics and time series 0.86
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.79
Abstract
Standard diagnostic procedures for assessing the goodness of fit of linear models include tests of the null hypothesis of no residual autocorrelation against the alternative of linear dependence. The literature proposes several portmanteau tests for residual autocorrelation in SARMA models, mainly focusing on two cases: short-term and single-seasonal autocorrelation. Since many time series exhibit multiple seasonal patterns, whose periodic components may interact with one another, diagnostic tests for residual autocorrelation in a multi-seasonal framework are needed. However, the literature still lacks portmanteau tests specifically designed for multiple seasonal autocorrelation. This paper addresses this gap by extending classical portmanteau tests to settings with multiple seasonalities. In addition, the proposed approach jointly tests for both short-term and multiple seasonal residual autocorrelation. Test statistics and their asymptotic distributions are defined. Then, Monte Carlo simulations are used to evaluate the performance of the tests and the consistency with the expected results is assessed using statistical tests. The results suggest that the proposed extension can serve as a useful goodness-of-fit diagnostic tool for the class of mSARIMA models. An application to road traffic data is also provided.
Naoki Hayashi, Yota Maeda, Yasushi Esaki
Tue 29 Sep
stat.MLcs.LGmath.STstat.TH
pdf html
- interest Kernels, operators and functional data 0.88
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.85
- reading log, theory, matrix factorization, true
Abstract
Non-negative matrix factorization (NMF) is a singular statistical model whose Bayesian asymptotics are governed by the real log canonical threshold (RLCT). We study the local geometry of the factorization map and derive an upper bound for the RLCT of NMF. Let $H$ be the model inner dimension and $H_0$ the non-negative rank of the true $M\times N$ matrix. Assuming that the true matrix admits a strictly positive factorization of inner dimension $H_0$ in the interior of the parameter domain, we prove, for smooth positive priors, that $\lambda\leq \{(H-H_0)\min(M,N)+H_0(M+N-H_0)\}/2$. This bound strictly improves the previous bound when $H_0\geq3$. The proof uses a local analytic normal form that separates independent linear coordinates from a residual matrix product. When $H=H_0$ also equals the ordinary rank of the true matrix, we obtain the exact value $\lambda=H_0(M+N-H_0)/2$. Under the standard assumptions of singular learning theory, these results bound the leading coefficients of the expected Bayesian generalization error and the Bayesian free energy.
Man Hei Ngou, Yanran Li, Zhexiao Lin, Zexi Cai
Thu 24 Sep
stat.ME
pdf html
- interest Causal inference and identification in econometrics 0.87
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
- favorites near “Difference-in-Differences with multiple time periods” 0.84
Abstract
Chatterjee's estimator uses only two local comparisons per interior observation, leaving a finite-replication variance gap under fixed alternatives. We introduce BRACE, short for blockwise rank aggregation for correlation estimation, which replaces adjacent comparisons with local blocks and averages every within-block rank difference. The block size controls the variance cost of finite local replication. An L2 expansion separates the efficient first-order component of the response-rank statistic from an orthogonal finite-replication component. For an admissible diverging block size, the latter vanishes, so the direct rank estimator attains the information bound. The decomposition yields a consistent variance estimator and Wald confidence intervals under fixed alternatives. At independence, the efficient first-order term vanishes, and the proposed estimator enters a second-order regime in which the same factor controls its null variance.
Hugo Gobato Souto
Wed 30 Sep
stat.MLcs.LG
pdf html
- reading mathematics, versus, causal, treatment
- your work near “Valid Inference with Imperfect Synthetic Data” 0.81
- favorites near “More power to you: Using machine learning to augment human coding for more efficient inference in text-based randomized trials” 0.87
Abstract
This paper proposes Wasserstein Causal Forests (WCF) for settings in which each unit's outcome is itself a probability distribution. This study also defines finite-grid transformed average and conditional average treatment effects, including a reference-distance contrast that asks whether treatment moves unit-level distributions toward a prespecified benchmark. Simulations cover null effects, location and shape changes, limited overlap, equal-mean but different laws, heterogeneous effects, multimodality, and structural zeros. WCF is most accurate on the conditional-law metric in most reported designs and sharply improves reference-effect estimation in the principal location-and-shape settings, but it is less accurate than the forest baselines for multimodal settings. WCF is applied to the famous Project STAR \citep{word1990state}, revealing that small classes alter more than the mean: they raise within-grade mathematics achievement by $0.161$ standard deviations on average (SE $0.028$); but the gain is not a uniform location shift, it is larger in the upper part of the classroom score distribution ($+0.179$ at the ninetieth percentile versus $+0.091$ at the tenth) and, in descriptive stratum estimates, largest in the schools serving the most economically disadvantaged students (highest free-lunch quartile, $+0.262$, versus $+0.092$ to $+0.164$ elsewhere), while overall dispersion is essentially unchanged.
Mark Patrick Roeling
Mon 28 Sep
cs.LGstat.COstat.ME
pdf html
- reading mathematics, empirical, functions, personal
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.83
- interest Computational methods for heterogeneous agent and dynamic economic models 0.79
Abstract
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable representation, spectrally diagnosed and pruned, and then certified exactly through multimodular evaluation. Exact certification makes the numerical proofs fully explicit and independently verifiable. This framework can be run on a standard personal computer. Across three extremal problems, we show that neural optimization can contribute to rigorous mathematics in three distinct ways: by discovering improved constructions, by exposing empirical invariants that lead to proofs, and by revealing optimization barriers whose geometry motivates new analytic or numerical representations. More broadly, these results suggest a path toward AI-assisted mathematics in which flexible computational discovery and exact certification become complementary components of a single rigorous workflow.
Haoyi Yang, Lingzhou Xue
Mon 28 Sep
math.STstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.89
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.87
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.84
Abstract
Combining $p$-values is a fundamental procedure in global hypothesis testing. In modern high-dimensional settings, however, component $p$-values often exhibit complex dependence and rely on asymptotic approximations rather than exact finite-sample uniform distributions. This paper establishes a unified asymptotic theory for weighted transformation statistics that decouples marginal finite-sample approximation error from the joint dependence structure. We also provide sufficient conditions based on conditional probability bounds to verify the joint-tail conditions. Utilizing this framework, we derive explicit dimension-growth and correlation rates for test statistics operating under asymptotic Gaussian and chi-square calibrations. For non-exact finite-sample statistics, we analyze standardized weighted sums, demonstrating how Cram\'er moderate deviations control relative tail error. Analytical examples demonstrate why both marginal and joint conditions are mathematically indispensable for valid global inference under dependence, and numerical experiments confirm that our asymptotic framework maintains accurate finite-sample size control at extreme significance levels.
Zhidi Lin, Yuhao Liu, Ying Li, Edwin Fong, Petar Djurić
Tue 29 Sep
cs.LGeess.SPstat.ML
pdf html
- reading state, equations, differential equations, irregular
- favorites near “Spectral State Space Models” 0.80
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.81
Abstract
Gaussian process stochastic differential equations (GP-SDEs) provide a flexible Bayesian model for unknown continuous-time state dynamics with uncertainty quantification, but learning and inference from noisy and irregular observations remain computationally challenging. To address this issue, we propose GP-SDE Matching, a simulation-free variational framework for Bayesian GP drift learning and continuous-time state smoothing. We analytically marginalize the sparse GP posterior to derive a tractable drift-matching objective that accounts for both the posterior mean and uncertainty of the unknown drift. To handle irregular observations, we further introduce an irregular-time-aware variational state posterior that incorporates the actual observation times during both encoding and continuous-time marginal querying. Experiments on the stochastic Lorenz--63 system demonstrate substantially improved drift recovery and state reconstruction under irregular observations, while five system identification benchmarks show robust forecasting under increasing observation sparsity and competitive performance against existing latent-SDE and state-space methods.
Dale S. Kim, Qing Zhou
Tue 29 Sep
stat.MEstat.ML
pdf html
- your work near “Automated Solution of Heterogeneous Agent Models” 0.86
- interest Semiparametric inference and debiased machine learning 0.87
- favorites near “Priors for the Long Run” 0.82
Abstract
Learning the structure of latent factor models involves two central challenges: (1) estimating the number of latent factors and (2) learning the support of the mapping from latent variables to observed variables. This is especially challenging for nonparametric regimes and nonlinear settings. We propose a method for latent structure learning, based on pairwise dependence measures on the observed variables using a graph-theoretic representation. We show that both the number of latent factors and nonlinear mapping structure can be identified from the distribution of observed variables under mild structural assumptions. Unlike prior work restricted to linear correlations, we establish identifiability and consistency for a general class of dependence measures under nonlinear factor models. This motivates a Dependence Thresholding (DT) algorithm, which jointly estimates the number of latent factors and nonlinear mapping structure from observational data alone. We pair this with a neural network architecture constrained by the nonlinear mapping structure, to recover the nonlinear function. Through simulation studies, we show that the DT algorithm is accurate in practice, even when using flexible methods such as neural networks, and exhibits robustness against violations of its assumptions in high-dimensional settings.
Thomas Dybdahl Ahle, Yaroslav Bulatov, Christopher De Sa, Christopher Ré
Wed 30 Sep
cs.LGstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.83
- favorites near “A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning” 0.88
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
Abstract
Recent work has established that power-law spectral conditions on data enable tight convergence bounds for deterministic gradient descent, resolving the conflict between classical exponential bounds and observed power-law learning curves. In this work, we extend this result to the stochastic regime of high-dimensional machine learning. We provide two main contributions: (1) We generalize the power-law spectral theory to Stochastic Gradient Descent (SGD), showing that the same spectral exponents govern stochastic dynamics; (2) For the fundamental case of isotropic Gaussian data, we provide a precise analysis of data shuffling, deriving exact constants that prove Single Shuffle is strictly superior to Flip-Flop and IID sampling. Our results bridge the gap between abstract spectral theory and practical stochastic training choices, offering a unified picture of how data geometry drives optimization speed.
Haoxuan Wang, Yuchen Fang, Sen Na
Mon 28 Sep
math.OCcs.LGstat.COstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.84
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.83
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
Abstract
We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challenges. Moreover, existing theoretical guarantees for constrained stochastic methods predominantly rely on bounded-variance assumptions, leaving the heavy-tailed noise regime largely unexplored. To address this gap, we propose a novel trust-region method within the stochastic sequential quadratic programming framework, termed TR-SSQP. Our method employs a normal-tangential decomposition in the step computation to balance optimality and feasibility. In addition, we incorporate a normalization mechanism in the design of the trust-region radius, together with Polyak momentum for gradient estimation, ensuring stable updates without gradient clipping. When the trust-region radius and the momentum parameter decay at appropriate rates, we establish global almost-sure convergence of the method. To the best of our knowledge, this is the first asymptotic convergence result for constrained stochastic optimization under heavy-tailed noise. We demonstrate the promising performance of the proposed method through extensive numerical experiments, including comparisons among its variants and with existing constrained stochastic optimization methods.
Tori Qiu, Ander Artola Velasco, Manuel Gomez-Rodriguez
Mon 28 Sep
cs.LGcs.AIstat.ML
pdf html
- reading mathematics, science, appear, generating
- favorites near “The Consensus Game: Language Model Generation via Equilibrium Search” 0.86
- interest Language models as statistical objects 0.83
Abstract
Self-consistency has become a popular technique for enhancing the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer through a majority vote. However, because model providers typically charge users in proportion to the number of reasoning paths generated, they have a financial incentive to artificially increase the path count. In this work, we show that an unfaithful provider can exploit this incentive using a simple, efficient algorithm while avoiding detection by an auditor: by generating and strategically reordering additional reasoning paths, the algorithm makes every path appear necessary to reach the majority. To validate our algorithm, we conduct experiments with multiple instruct models from the Llama and Qwen families, as well as reasoning models distilled from DeepSeek-R1, on benchmark datasets spanning mathematics, science, and question answering. Our results suggest that the distribution of additional reasoning paths generated by our algorithm is heavy-tailed and that substantial capacity to overcharge remains even under the best possible audit designed to keep the false-positive rate below $\alpha = 0.1$.
Dae Woong Ham, Xuejun Zhao, Stefanus Jasin, Fenghua Yang
Fri 25 Sep
cs.AIcs.ITmath.ITstat.ME
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.82
- interest Adaptive data collection and anytime-valid inference 0.84
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.80
Abstract
Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, the decision maker can selectively query AI, send an item directly to a human, escalate an AI-scored item to a human after observing the AI report, or stop once sufficient evidence has accumulated. We derive an information-theoretic lower bound that captures the minimum cost of achieving prescribed testing errors and characterizes the value of AI information and human verification through a report-dependent information frontier. Motivated by this characterization, we develop SCALE, a sequential cost-aware policy that combines selective AI scoring with adaptive human escalation. SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. We further extend the framework to an unknown AI-output model using paired AI-human pilot data. Numerically, SCALE approaches Human-only or AI-only testing when one source clearly dominates, while achieving its largest savings when inexpensive AI judgments and selective human verification are both valuable.
Louis Yiven Zhu
Fri 25 Sep
cs.LGcs.CLstat.ME
pdf html
- interest Semiparametric inference and debiased machine learning 0.85
- reading bank, theory, measured, makes
- your work near “Valid Inference with Imperfect Synthetic Data” 0.80
Abstract
An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that three widely portable ones mean something different for a judge than for a test because the judge setting rearranges the roles those designs rest on. First, an internal-consistency coefficient computed over rubric elements contains no scorer facet. Holding one judge's measured error rate fixed at $4.72\%$, KR-20 still ranges from $0.01$ to $0.68$ as the item bank is redesigned around it, and varying judge error moves the coefficient by a comparable amount, so item design and judge error are not separately identified and no single value can be read as a property of the judge. Second, the dependability index $\Phi(\lambda)$ is a ratio of variance components, and the classification probability with which it is sometimes identified differs from it by $0.25$-$0.43$ on our bank and by $0.17$-$0.30$ on simulated data where the underlying model holds exactly. Third, Livingston-Lewis accuracy is indexed to an examinee's own true score on the same instrument, so scoring it against external gold conflates judge unreliability with criterion invalidity. Reviewing the three closest judge-evaluation papers, we found no published instance of these errors, which makes the caution prospective. A coefficient that cannot be attributed to the judge nonetheless travels downstream into deployment decisions and disclosure documents. We therefore close with four reporting lines that keep the attribution attached to the number.
Wei Biao Wu
Mon 28 Sep
stat.MLcs.LG
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.81
- reading argument, inequality, moments, ordinary
- favorites near “Majorizing Measures, Sequential Complexities, and Online Learning” 0.86
Abstract
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displacement argument yields $T^{-1/3}$ expected average squared-gradient stationarity with a horizon-dependent stepsize. An explicit oracle-complexity corollary matches the known smooth Blum--Gladyshev (BG-0) lower bound, including the $Lb_2\Delta^3\varepsilon^{-6}$ and $L\Delta\sigma^2\varepsilon^{-4}$ stochastic terms, where $\Delta$ is the initial objective gap and $\sigma^2+b_2\|x-x_1\|^2$ bounds the variance. Thus unchanged SGD attains the minimax stochastic complexity in this second-moment class. For $p>2$, predictable localization and a Hilbert-space Fuk--Nagaev inequality yield a high-probability bound separating logarithmic variance and polynomial rare-shock contributions. The localization radius is derived from the recursion: no bounded-iterate assumption, clipping, normalization, momentum, or increasing batch size is needed. We also give increasing-confidence rates, an objective-gap-growth refinement recovering root-$T$ stationarity, and stochastic $L^p$-Lipschitz examples. The broad BG-0 optimality statement is distinguished from the smaller mean-square-smooth class, in which additional oracle structure permits faster algorithms.
Christopher A. Lindley, Nikolaos Dervilis, Keith Worden
Mon 28 Sep
stat.MLcs.LGcs.SYeess.SYstat.CO
pdf html
- reading equation, mcmc, dynamical, structure-preserving
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.81
- interest Computational methods for heterogeneous agent and dynamic economic models 0.78
Abstract
Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based identifiers rely on evolutionary optimisation and return point estimates of the model structure. This paper instead proposes the TAG framework within a Bayesian setting. A generative prior is defined over tree structures and their parameters, and a Reversible-Jump MCMC sampler with structure-preserving tree moves is used to infer the joint posterior over model structure, parameters and predictions. Two training objectives are considered; that is, a one-step-ahead objective with conjugate parameter proposals, and a simulation-based objective handled by likelihood-free inference. The approach is validated on a simulated polynomial NARX system, the Silverbox benchmark, and wave-loading data from the Christchurch Bay Tower, where embedding Morison's equation as a fixed initial tree yields a grey-box model that outperforms the physics-driven baseline. The results demonstrate that Bayesian TAGs are well suited to quantifying uncertainty in equation discovery for dynamical systems and to fitting physics-informed models.
Paul T. von Hippel
Wed 30 Sep
stat.ME
pdf html
- reading empirical, strongly correlated, estimates, implementations
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.79
- favorites near “Labor Rationing” 0.76
Abstract
Education agencies often summarize test score distributions by counting how many students scored in 3 to 5 different \textit{bins}. The HETOP model transforms bin counts into estimated means and standard deviations, assuming that scores follow a normal distribution within each school or district. Past HETOP implementations ran slowly, taking 3--60 minutes, if they finished, when given bin counts for all 1,151 districts in Texas. Our new function, \code{fast\_hetop()} in the R package \pkg{binest}, runs all Texas districts in less than a second. Users can choose between maximum likelihood, empirical Bayes, sample-based or population-based estimates. Estimates are strongly correlated with true values, but have bias when the score distribution is skewed and scores are concentrated in the lowest or highest bin.
Niklas Koenen, Claudia Battistin, Jeriek Van den Abeele, Martin Jullum
Tue 29 Sep
stat.MLcs.LG
pdf html
- interest Semiparametric inference and debiased machine learning 0.88
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.82
- favorites near “When do common time series estimands have nonparametric causal meaning?” 0.82
Abstract
Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.
Noa Rubin, Zohar Ringel
Fri 25 Sep
cs.LGcond-mat.dis-nnstat.ML
pdf html
- reading theory, regulator, iterated, algorithmic
- favorites near “Finding Regularized Competitive Equilibria of Heterogeneous Agent Macroeconomic Models via Reinforcement Learning” 0.80
- your work near “Timing as an Action: Learning When to Observe and Act” 0.78
Abstract
Despite the importance of reinforcement learning with verifiable rewards (RLVR), the extent to which it can learn new reasoning capabilities remains debated. Here we study the optimization landscape of RLVR on algorithmic tasks, such as iterated group and quasigroup multiplication. To this end, we map entropy-regularized RLVR over myopic tabular policies onto an energy-based (spin-glass) model over deterministic policies. This mapping upper-bounds what RLVR can achieve, and lets us rigorously characterize the landscape in this tabular setting. We show, both theoretically and experimentally, that for a wide class of models and tasks with uncorrelated inputs, this landscape is benign, containing no local minima that could trap RLVR training. Rather, the practical difficulty of these tasks appears to stem, at least in part, from issues such as diffusive barriers and gradient-estimation error in traversing the landscape. These are genuine obstacles that can prevent a solution from being found, but they are distinct from the landscape itself being rugged. We show that these obstacles can often be mitigated through the choice of entropy regulator. Consistent with this theory, we find that a transformer trained from scratch, using only last-token rewards, successfully learns an algorithmic chain of thought for iterated non-Abelian group multiplications.
Xuan Liu, Jingbin Qian, Haosheng Chen
Wed 30 Sep
cs.LGstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.87
- reading gains, policy, recovered, large
- interest Language models as statistical objects 0.82
Abstract
Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR). An exact verifier can also support test-time scaling by selecting a passing response from multiple samples, while other deployments use beam search, adaptive sampling, or tools. We study single-sample decoding, where each query receives one response without search, to ask whether search-exposed behavior can be absorbed into the model. Existing verified-response post-training recipes do not generally distinguish problems already solved on the first decode from failures recovered within K samples. Under a fixed budget, this can spend examples repeating behavior the deployed policy already has. We introduce GapFT, which selects training evidence by the source checkpoint's single-sample outcome and fine-tunes on the Pass@K-Pass@1 gap: problems the policy fails on one sample but solves within K samples. We match training examples, processed tokens, and optimizer steps while keeping the objective unchanged. GapFT fills the matched budget with recovered failures and uses an exact decomposition to distinguish corrections of recovered and missed failures from regressions on first-decode successes. On LogiQA 2.0 and ReClor with Llama-3.1-8B, GapFT improves Pass@1 by 14.4 and 13.9 points over the source model, outperforms budget-matched uniform verified RFT at the same learning rate, and matches fine-tuning on the full verified pool using one third of the data. A single decode matches the source model's verifier-selected Pass@4 accuracy. A randomized control attributes gains to covering distinct failures, and our analysis relates available gains to transferable failure support. A three-seed Qwen2.5-7B replication retains positive gains over uniform RFT on both logic tasks.
Ariel Linden
Thu 24 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.87
- reading adjustment, nominal, treatment, basis
- interest Causal inference and identification in econometrics 0.81
Abstract
Studies enrolling participants on the basis of an extreme baseline value are susceptible to regression to the mean (RTM), such that some observed pre-post change is expected even without treatment. Existing methods estimate and decompose RTM retrospectively. We extend this framework to prospective study design by deriving a closed-form sample-size method for continuous, cutoff-selected, single-arm pre-post studies. The method partitions anticipated total change into that expected from RTM and the residual treatment effect, and powers the study to detect the latter. It uses the general bivariate-normal RTM expression, allowing baseline and follow-up variances to differ, and incorporates the corresponding conditional variance of the pre-post change. To our knowledge, no published method or software implements this cutoff-based RTM framework for prospective closed-form sample-size determination in this setting. The Stata command power onemean_rtm provides solutions for sample size and minimum detectable effect, normal-theory power evaluation, attrition adjustment, and sensitivity analysis. Monte Carlo simulation across 58 scenarios demonstrated accurate Type I error control and showed that achieved power rapidly approached nominal power as sample size increased, with the closed-form calculation generally conservative at very small sample sizes.
Aaditya Ramdas
Tue 29 Sep
math.STcs.GTcs.ITmath.ITstat.MEstat.TH
pdf html
- reading wealth, capital, log, actually
- interest Adaptive data collection and anytime-valid inference 0.85
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.81
Abstract
Bell and Cover showed that an investor who multiplies the initial unit of capital by an independent uniform random variable on $(0,2)$, and then uses the log-optimal portfolio, wins a head-to-head wealth comparison with probability at least one half against every independently randomized competitor. We explain very simply how this result transfers to testing by betting: for any composite null $\mathcal P$ and simple alternative $Q$, denoting $E^*$ as the corresponding numeraire e-variable, we show that $UE^*$ exceeds any other e-variable $E$ with probability at least half. Interestingly, we show that this competitive optimality result is actually equivalent to the numeraire inequality $\mathbb E_Q[E/E^*]\leq1$, and in general randomization only helps the numeraire and fails to improve the competitive advantage of an arbitrary e-variable. Under optional stopping with or without knowledge of $U$, we emphasize a key distinction between e-process validity and competitive optimality. We also show that competitive optimality comes at the price of expected log wealth and power: thresholding $UE^*$ at $1/\alpha$ has sharp size at most $\alpha/2$, but the factor of two actually disappears under optional stopping. Even after correcting for this factor of two, the test is dominated in conditional rejection probability by randomizing the testing threshold (randomized Markov's inequality). Thus, Bell-Cover randomization is optimal for a specific competitive objective, at the cost of others.
Yingbo Zhao, Zeyu Yang, Zhoufan Zhu
Tue 29 Sep
stat.MLcs.AIcs.LG
pdf html
- favorites near “Finding Regularized Competitive Equilibria of Heterogeneous Agent Macroeconomic Models via Reinforcement Learning” 0.84
- interest Computational methods for heterogeneous agent and dynamic economic models 0.79
- your work near “Timing as an Action: Learning When to Observe and Act” 0.79
Abstract
Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.
Renato Valladares Panaro, Vinícius Mayrink, Fábio Demarqui
Tue 29 Sep
stat.ME
pdf html
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.88
- reading functions, likelihood bayesian, familiar, po
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.84
- looks like clinical and biomedical studies
Abstract
We present spsurv, an R package for semi-parametric time-to-event regression based on Bernstein-polynomial estimation of unknown baseline functions. The package provides a unified modelling interface for proportional hazards (PH), proportional odds (PO), and accelerated failure time (AFT) models for right-censored data, with either maximum likelihood or Bayesian estimation via Stan. Smooth baseline hazard, odds-function, or log-time structures are estimated without assuming a parametric baseline family, while retaining familiar hazard-ratio, odds-ratio, and time-ratio interpretations. We describe methodology, implementation, and syntax; evaluate finite-sample behaviour in a Monte Carlo study; and illustrate usage with oncology trials.
David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin
Wed 30 Sep
cs.AIstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.83
- interest Language models as statistical objects 0.84
- reading tell, language, language model, explains
Abstract
Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer's credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer's width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.
Jiayu Li
Fri 25 Sep
cs.LGq-fin.STstat.ML
pdf html
- reading future, causal, expanding, going
- interest Macroeconometrics and time series 0.84
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.79
Abstract
Validating a model on a time series asks for three things at once: each training run should use most of the sample (sufficiency), the test sets should together cover most of the sample (coverage), and training data should come before test data (causality). We prove that the three cannot be had together and price each one. Let $\alpha$ be the smallest training fraction over folds, $\beta$ the fraction of the sample covered by tests, $\Lambda$ the fraction of the sample used as training data from the future of a test point, and $\delta$ the distance from a test point to the nearest training point in its future. Every scheme on a sample of length $T$ satisfies $\alpha+\beta \le 1+\Lambda$ and $\alpha+\min\{\beta,\delta/T\} \le 1$, and under $\beta$-mixing the leakage bias at a test point is at most $2M\beta_{\mathrm{mix}}(\delta)$. In words: going beyond the causal frontier $\alpha+\beta=1$ requires training on the future; that future data must sit within $(1-\alpha)T$ of a test point; and its harm depends on its distance, not its amount. Hence expanding walk-forward is exactly the Pareto frontier of causal validation, $k$-fold cross-validation buys the most future data, and purged $k$-fold with an embargo pays in distance instead, which is cheap when the process forgets quickly but cannot repair the part of causality demanded by non-stationarity. On pure noise, shuffled 5-fold reports an information coefficient of $+0.32$, while contiguous 5-fold, using the same amount of future data, reports $+0.004$.
Chen Cheng, Ruiting Liang, Rina Foygel Barber
Thu 24 Sep
math.STcs.LGstat.MEstat.MLstat.TH
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.82
- interest Language models as statistical objects 0.84
- your work near “Timing as an Action: Learning When to Observe and Act” 0.79
Abstract
We introduce Rolling Conformal Prediction (rolling-CP), a distribution-free predictive inference method for the setting of sequential model training. Specifically, given a data stream $(X_1,Y_1),(X_2,Y_2),\dots$, at each time $n$ the trained model may depend on the observed history $\{(X_i,Y_i)\}_{i<n}$. This setting arises naturally in modern sequential training, including one-pass training over massive datasets and continual fine-tuning or test-time adaptation of language models during deployment. Rolling-CP first calibrates each incoming observation against the current predictor and then rolls it into future training. In this way, we avoid the need for data splitting. Remarkably, although the models at times $n=1,2,\dots$ may have entirely different properties and accuracy levels, for exchangeable data it is nonetheless possible to establish a guarantee of marginal coverage, with a familiar universal factor-two guarantee (a worst case guarantee of $1-2\alpha$ coverage, as compared to the target level $1-\alpha$), without any assumptions of stability or any restrictions on the model training process. For i.i.d. data streams, we further prove high-probability training-conditional validity uniformly over time; under stability conditions, coverage guarantees sharpen towards $1-\alpha$. Numerical experiments on sequential regression, multiclass SGD, and one-pass neural-network training further demonstrate the practical effectiveness of rolling-CP.
Jeroen P. Jansen, Harlan Campbell, Shannon Cope, David M Phillippo, Antonio Remiro-Azócar
Tue 29 Sep
stat.MEstat.AP
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
- reading versus, treatment, document, wrong
- interest Causal inference and identification in econometrics 0.81
Abstract
Objectives: There has been increased appreciation of the differences between marginal and conditional estimates and different types of effect measures regarding their applicability to different target populations. This issue of transportability is of concern in model-based cost-effectiveness analysis (CEA) when treatment effects from (international) trials are applied to (country-specific) baseline risk estimates. The objective of this paper is to create awareness regarding the issues that arise when using different types of treatment effect and baseline risk estimates in a model-based CEA to inform health technology assessment (HTA). Methods: We clarify collapsibility, marginal versus conditional estimation, and transportability; derive the ideal modeling approach implied by a marginal cost-effectiveness estimand; and characterize the issues of common modeling approaches, illustrated with a fictitious state-transition model. Results: An individual-level simulation that predicts outcomes from conditional inputs and averages them over the target population targets the marginal cost-effectiveness estimand. Cohort-model approaches that marginalize inputs early, evaluate an outcome regression model at mean covariates, or combine a conditional effect with a marginal baseline (or vice versa) can misstate cost-effectiveness results even with correct-population inputs; inputs from the wrong population add further error. Conclusions: The most rigorous approach is an individual-level simulation that carries conditional inputs (baseline risk, treatment effect, prognostic effects and effect modifiers) and marginalizes late. Cohort approaches instead marginalize early, relying on aggregated inputs, and do not necessarily target the marginal cost-effectiveness estimand of interest for HTA. Model developers should document, for each input, whether it is marginal or conditional and its population.
Ziang Yuan, Xinran Li, Shuangning Li
Tue 29 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.90
- interest Causal inference and identification in econometrics 0.83
- favorites near “Policy Learning with Observational Data” 0.88
Abstract
Covariates are widely used in randomized experiments to improve precision. However, in the presence of interference, where outcomes may depend on the treatment assignments of other units, standard covariate adjustment methods may fail to preserve desirable properties such as the no-harm property, meaning that incorporating covariates does not worsen estimator performance. Existing approaches that incorporate covariates under interference typically rely on specifying a particular interference structure, and their guarantees can be sensitive to misspecification. In this paper, we study how to incorporate covariate information in a way that preserves the no-harm property while remaining largely agnostic to the underlying interference structure. We focus on the estimation of the expected average treatment effect (EATE) using the H\'ajek estimator and study rerandomization under interference, a design-stage procedure that restricts the assignment space to allocations with sufficiently small covariate imbalance. Our framework also allows the covariates used for rerandomization to depend on the treatment assignment itself, such as the proportion of treated neighbors, which naturally arises in settings with interference. We show that, even under interference, rerandomization can improve estimation precision asymptotically relative to unrestricted Bernoulli randomization, relying only on mild conditions on the dependence structure across units. When a conservative dependence graph is available, we further develop an optimization-based conservative variance estimator for inference under rerandomization.
Tobia Filosi, Emilio Porcu, Claudio Agostinelli
Tue 29 Sep
math.PRstat.ME
pdf html
- reading processes, functions, variational, natural
- interest Kernels, operators and functional data 0.83
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.83
Abstract
Metric graphs are generalisations of linear networks and provide a natural framework for the definition of continuously-indexed Gaussian processes. We define a new class of distances on these topologies, termed polyharmonic distances, which unify and extend the spectral construction underlying the effective resistance distance and the biharmonic one. We give both a spectral and a variational characterisation. Furthermore, we show an explicit class of stochastic processes whose variograms coincide with the squared polyharmonic distances. Finally, we show how these metrics can be composed with suitable completely monotonic functions to define isotropic processes having any prescribed finite-order mean-square differentiability along the edges and satisfying the Kirchhoff conditions up to order one at the vertices.
Percy S. Zhai, Maggie Cheng, Wei Biao Wu
Tue 29 Sep
math.STstat.MLstat.TH
pdf html
- reading long memory, series, processes, theory
- interest Macroeconometrics and time series 0.83
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.78
Abstract
We study conformalized quantile regression and conformalized median regression that fit a model on one block of a time series and calibrate the conformal interval on the adjacent block. The existing theory of conformal prediction for time series rests largely on mixing conditions, which are hard to verify from a time-series model and fail for many standard processes, including simple ones with short memory. We replace this theoretical toolbox with the functional dependence measure, which in principle accommodates long-memory observations. The accuracy of the conformal interval length for time series has been understudied. To the best of our knowledge, this paper is the first work that establishes non-asymptotic coverage guarantees and accuracy of interval length simultaneously for split conformal regression on time series. Furthermore, for Gaussian linear processes with long memory, where both the estimation of the center and its calibration converge slowly, we establish a sharper rate for the length error. We show that the calibrated length converges faster than the estimated center itself, and provide a matching lower bound for the usual centers when the calibration block is sufficiently large relative to the training block. To our knowledge, this is the first theoretical analysis of conformal interval length dedicated to long memory.
Jinyuan Chang, Jing He, Weiming Li, Chen Lin
Tue 29 Sep
math.STecon.EMstat.TH
pdf html
- interest Macroeconometrics and time series 0.87
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.82
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.78
Abstract
We propose a new $L_2$-type test for white noise which allows the dimension $p$ of the time series to either (i) be a fixed constant, or (ii) diverge with the sample size $n$. The proposed test statistic exhibits an interesting phase transition, following two different regimes of behavior: $p$ is fixed, and $p\rightarrow\infty$. Because identification of the operable regime is difficult, if not impossible in practice, we devise a novel adaptive bootstrap method to construct unified testing procedure across different phases. Numerical experiments confirm the good finite sample performance of the proposed adaptive $L_2$-type test in comparison to the existing methods in the literature. The proposed testing procedure has been implemented in R package HDTSA.
Prasanjit Dubey, Xiaoming Huo
Thu 24 Sep
stat.MEmath.STstat.TH
pdf html
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.88
- interest Adaptive data collection and anytime-valid inference 0.87
- favorites near “Majorizing Measures, Sequential Complexities, and Online Learning” 0.84
Abstract
We study sequential testing of a fixed family of hypotheses when observations are costly and a sampling horizon is specified in advance. The challenge is to pool evidence for earlier decisions when the number and identities of false hypotheses are unknown, while controlling the probability of any false rejection at level $\alpha$. Existing merges attain the pooled growth rate at a single number of false hypotheses: averaging when one is false, multiplying when all are. Under an independent-stream model with common simple null and alternative distributions, we test each intersection with a prior-weighted mixture of products of marginal likelihood ratios. Closed testing combines these elementary-symmetric-polynomial mixtures to identify individual false hypotheses. Design-specific boundary calibration gives finite-horizon family-wise error control, with exact finite-state guarantees or a confidence qualification for Monte Carlo calibration. The prior-matched mixture uniquely maximizes expected log evidence at each horizon. Mixtures assigning positive weight to every nonempty subset of streams attain log-growth rate $lD$ when $l$ streams follow the alternative. Here $D$ is the mean log likelihood ratio per alternative observation, and a round supplies one observation per stream. This rate attains the first-order intersection-delay lower bound as $\alpha\downarrow0$ at fixed dimension, configuration, weights, and a long enough horizon. Power for an individual hypothesis cannot exceed the best single-stream power at a given deadline, but closure removes the multiplicity penalty when all are false. Gaussian, basket-trial, language-model and advertising studies illustrate both. The primary basket boundaries are 40-52% below $1/\alpha$. Across 41 simulated configurations, the median reduction in capped mean patient outcomes relative to prespecified interim-look Bonferroni tests is 31%.
Qianyi Chen, Bo Li
Tue 29 Sep
cs.LGstat.ME
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.85
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.86
- interest Language models as statistical objects 0.83
Abstract
Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or general-purpose language models. We study whether these synthetic labels can improve conditional coverage when only a small trusted sample is available. Building on score-quantile regression, we introduce prediction-powered quantile learning: a synthetic-labeled pool estimates pinball risk, paired trusted and synthetic outcomes correct its bias, and an independent trusted split performs final conformalization. Profiling pinball risk over scalar corrections reveals that population conditional-coverage error is its functional gradient; the corresponding Hessian removes global shifts and weights remaining shape error by boundary density. Composing this geometry with prediction-powered learning yields a three-resource expansion and a benefit--cost rule for synthetic power. Across eight regression benchmarks, synthetic-powered quantile learning substantially improves downstream conditional coverage while preserving marginal validity and producing more compact prediction sets. A human-rating study finds similar gains from external LLM labels and exposes a quality--quantity--cost tradeoff.
Ben Adcock, Simone Brugiapaglia, Xuemeng Wang
Wed 30 Sep
math.NAcs.LGcs.NAstat.ML
pdf html
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.84
- interest Kernels, operators and functional data 0.86
- reading operators, operator, functions, theory
Abstract
Out-of-distribution (OOD) generalization is a central challenge in scientific machine learning. We study regression problems in which the test distribution differs from the training distribution and ask: under what assumptions on the target function or operator is stable extrapolation possible, and how far beyond the training domain can one extrapolate? Existing theory controls the test error through additive penalties measuring the discrepancy between the training and test distributions. Such guarantees show robustness to small distribution shifts, but can very pessimistic in comparison to OOD performance observed empirically. We identify classes of holomorphic functions and operators for which the OOD generalization error converges at algebraic rates even in the presence of large distribution shifts. This phenomenon stems from the increasing smoothness of higher-index coordinates, leading to what we term a `blessing of high dimensionality'. For learning with either polynomials, deep neural networks or deep neural operators, we derive explicit rates for arbitrary test measures supported on suitable domains and quantify how the admissible domain depends on the underlying regularity of the function or operator. Our extrapolation guarantees are independent of the test distribution, depending only on its support. We also present a series of numerical experiments across a range of functions and operators that support the main theoretical findings.
Roberto Bruno, Adrien Vandenbroucque, Amedeo Roberto Esposito
Thu 24 Sep
cs.ITmath.ITmath.STstat.TH
pdf html
- favorites near “Orthogonal Statistical Learning” 0.90
- interest Adaptive data collection and anytime-valid inference 0.85
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.81
Abstract
We study asymmetric simple binary hypothesis testing between $H_0:P_0^{n}$ and $H_1:P_1^{n}$, based on $n$ independent and identically distributed observations. Leveraging a variational representation of R\'enyi divergence of order $\alpha$, we derive our main result: a finite-sample converse with $\alpha>1$. The bound uses both directions of the divergence $D_\alpha(P_1\|P_0)$ and $D_\alpha(P_0\|P_1)$, tensorises under product measures, and contains familiar data-processing converses as boundary cases. For comparison, we apply the same variational approach to general $f$-divergences and specialise it to total variation, $E_\gamma$, Hellinger, and Kullback Leibler divergences, thereby recovering familiar converses within a unified framework. Together with an achievability bound involving R\'enyi divergence with $\alpha\in (0,1)$, the main converse recovers the phase transition of the optimal Type II error under the exponentially decaying Type I error constraint $\varepsilon_n=e^{-nr}$. Under regularity conditions, the optimal Type II error vanishes exponentially when $r<D(P_1\|P_0)$ and converges exponentially fast to one when $r>D(P_1\|P_0)$. We also derive sample-complexity bounds and extend both the converse and achievability analyses to locally differentially private observations, quantifying the cost of privacy and recovering the non-private achievability bound as the privacy constraint vanishes.
Eduardo Janotti, Lígia Henriques-Rodrigues, Antonio Carlos Pedroso de Lima
Thu 24 Sep
stat.MEmath.STstat.APstat.TH
pdf html
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.78
- interest Semiparametric inference and debiased machine learning 0.85
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.85
Abstract
We propose a class of extended generalized Pareto models for right-censored survival data, with particular emphasis on tail inference and long-term extrapolation. Our framework integrates extreme value theory and survival analysis, combining a generalized Pareto distribution with a flexible perturbation distribution on the unit interval. We consider three perturbation specifications: a parametric Beta model, a Bernstein polynomial estimator, and a histogram-based estimator. To facilitate direct comparison, all three models are fitted using a unified iterative procedure adapted to right censoring through Kaplan-Meier-based pseudo-observations. A Monte Carlo study evaluates finite-sample performance across different tail indices, censoring levels, sample sizes, and model complexities. The results reveal a trade-off between flexibility and stability: the Beta specification generally performs best for tail-index estimation, whereas the histogram estimator performs particularly well for scale estimation under low censoring. The Bernstein estimator shows intermediate performance and greater sensitivity to sample size and censoring. Applications to bladder cancer recurrence and heart-failure survival data show that models with very similar in-sample fits can nevertheless produce markedly different tail-index estimates and long-term extrapolations. These findings emphasize the importance of perturbation specification when extended generalized Pareto models are used for survival extrapolation under censoring.
Gordei Verbii
Fri 25 Sep
stat.MLcs.LG
pdf html
- interest Adaptive data collection and anytime-valid inference 0.86
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- favorites near “The Minimax Regret of Sequential Probability Assignment, Contextual Shtarkov Sums, and Contextual Normalized Maximum Likelihood” 0.86
Abstract
Worst-case Optimal Recovery (OR) and Bayesian learning describe the same Gaussian-quadratic-Hilbert problems in two vocabularies. We sharpen the correspondence - the radius of information equals a nugget-optimized GP posterior variance and is attained by the posterior mean at a closed-form balance nugget - and measure, inside three published Bayesian systems, where the worst-case side pays. The ledger is two-sided: the losses instruct as much as the wins. Morozov calibration tracks a test-access oracle within $1.00$-$1.19\times$ where $\sigma$-blind rules fail, is $4.9$-$6.3\times$ more reproducible across noise draws ($p=0.002$-$0.004$), and is the only deployable rule whose selection survives a change of backend ($1.36\times$ against $12$-$30\times$ for the released weight, ML-II and GCV); tight certificates cover at the information-theoretic floor with no numerical slack. But on exchangeable data split-conformal beats the OR head on interval score, a water-filling prior adds nothing without an oracle noise hint, and under covariate shift the OR band keeps coverage on every dataset yet loses interval score to split-conformal, and to a feature-free constant band, on most cells; what pays is not shift but shift on a learnable target, which a training-free audit statistic predicts before any model is fitted. In Bayesian optimization the certified width is a validity floor whose scalar inflation we prove inert under a checkable margin condition and check at every step. Inertness is graded, not binary, and in the size of the inflation as much as in the objective: $\kappa{=}2$ is inert wherever $\kappa{=}5$ is and on more cells besides, while $\kappa{=}5$ moves half the Ackley seeds and every Griewank seed. Exploration is a shape problem, not a scale one. The design rule: match the guarantee tool to the data regime, and audit the regime first.
Steve Lawford
Fri 25 Sep
math.STstat.MEstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.85
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.85
- reading govern, root, density, uniformly
Abstract
Under Gaussian sampling, removing location and scale turns the sample into a direction that is uniformly distributed on a sphere of dimension $n-2$, and sample skewness and kurtosis become a cubic-quartic polynomial image of spherical measure. We use this geometry to obtain exact finite-sample results for the Jarque-Bera statistic $\mathrm{JB}_n$. First, for every $n\ge 3$ the largest attainable value of $\mathrm{JB}_n$ is $n\{(n-2)^4+4(n-1)^2\}/\{24(n-1)^2\}$, attained by a single-outlier sample, so the asymptotic $10\%$ and $5\%$ tests have size zero for $n\le 6$ and the $1\%$ test for $n\le 7$; near this maximum the density of $\mathrm{JB}_n$ behaves like an explicit multiple of $(J_n^+-x)^{(n-4)/2}$. Second, passing from residual coordinates to power sums writes the joint density of skewness and kurtosis as an integral of the reciprocal square root of a polynomial discriminant; for every $n\ge 5$ the innermost integral runs over a single interval and is a Lauricella $F_D$ period. Third, this yields explicit laws: an arcsine law for $n=3$, an algebraic joint density for $n=4$, and a single complete elliptic integral, equivalently ${}_2F_1(\tfrac12,\tfrac12;1;\cdot)$, for $n=5$. Residual-coordinate collisions generate the discriminant singularities of the joint law, while stationary points of $\mathrm{JB}_n$ on the residual sphere govern the singularities of its one-dimensional density.
Chloé Hashimoto-Cullen, Amaury Durand, Laurent Bozzi, Benjamin Guedj, Yannig Goude, Sylvain Le Corff
Wed 30 Sep
stat.MLcs.LGstat.ME
pdf html
- favorites near “Spectral State Space Models” 0.80
- reading series, forecasting, frequencies, foundation models
- interest Macroeconometrics and time series 0.82
Abstract
Time Series Foundation Models (TSFMs) currently provide state-of-the-art results in forecasting tasks. They are available out-of-the-box and rely on in-context learning to make their predictions, which makes the quality of their performance highly sensitive to the user-selected lookback, covariates, horizon and training data distributions. In practise, the quality of the forecasts are variable but complementary, which highlights the need for a principled ensembling approach, rather than selecting the best context. This paper introduces Latent Inference-Time Guidance for TSFMs, which adaptively combines a pool of TSFM forecasts through a time-dependent latent space with independent components. The framework comes equipped with identifiability and reconstruction guarantees, whilst maintaining the off-the-shelf aspect of foundation models. We provide experiments on datasets at various frequencies and from multiple domains: these show that the approach is competitive with traditional ensembling approaches.
Philip Dawid
Fri 25 Sep
math.STstat.TH
pdf html
- reading presents, rules, possibly, process
- interest Adaptive data collection and anytime-valid inference 0.84
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.81
Abstract
This note presents a general form for a proper scoring rule that can be used for possibly right-censored survival data without requiring any knowledge or assumptions about the censoring process.
Marthe Elisabeth Aastveit, Alex Lenkoski, Thordis Thorarinsdottir
Mon 28 Sep
stat.MEstat.AP
pdf html
- reading markets, forecasting, curves, partially
- favorites near “Measuring Price Selection in Microdata: It's Not There” 0.81
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.81
Abstract
This paper develops prediction models for population-level survival curves observed over time and sampled from a heterogeneous mix of populations. We consider a discrete-time setting where each curve is only partially observed and forecasts of the remaining trajectory are needed for downstream decision making. Our approach recasts cross-population heterogeneity into a multivariate sampling model. We propose two forecasting models for partially observed curves: a full factor analysis model that extends a general factor representation to incorporate the partially observed survival curve, and a double PCA model. The methodology is motivated by demand forecasting in short-term rental markets, where market-level occupancy paths can be viewed as survival curves over the booking horizon and where forecasts of future occupancy feed into dynamic pricing algorithms. We apply the models to the newly released Wheelhouse dataset, which contains time series of market occupancy curves for 500 markets from 2017 to 2022. Model performance is assessed using the integrated quadratic distance, and we compare the proposed PCA-based methods to Holt's linear trend model across multiple forecast horizons. The results show that the proposed models yield accurate and stable forecasts of the remaining survival trajectory and generally outperform Holt's method, particularly at longer horizons.
Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
Mon 28 Sep
cs.AIstat.APstat.MEstat.ML
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.81
- interest Empirical Bayes, shrinkage and decision theory 0.86
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.84
Abstract
Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human preference target. Theoretically, we study three aspects of DIAL: (i) identification of latent LLM preferences, position effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against limited human evidence; and (iii) fixed-weight uncertainty quantification for the calibrated human preference. Empirically, we evaluate position debiasing and human alignment separately in controlled simulations and on three human-preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human-aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. Our real-data study collects over 410K judgments from 21 LLM judges in both display orders, providing a resource for future studies of LLM-judge bias, heterogeneity, and human alignment.
Leonhard Grosse, Sara Saeidian, Tobias J. Oechtering, Mikael Skoglund
Thu 24 Sep
cs.ITcs.CRmath.ITmath.STstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.84
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.81
- reading theory, control, uniformly, large
Abstract
We investigate $c$-interior pointwise maximal leakage (PML) as a tool for contraction analyses and disclosure control. Based on the strong adversarial threat models from maximal leakage, $c$-interior PML generalizes local differential privacy (LDP) to data-generating distributions with densities uniformly bounded away from zero by $c>0$. Viewing $c$-interior PML as an algebraic constraint on a kernel yields more flexible (and often tighter) contraction analyses than standard LDP. We provide tight bounds on the Dobrushin coefficient, and bound the contraction coefficient of the Hockeystick-divergence. We further derive strong data processing inequalities on $f$-divergences under $c$-interior PML constraints when the input distributions to the divergence are restricted to be in the $c$-interior. These results extend beyond the regime of pure LDP to cover a larger class of kernels, including, e.g., arbitrary stochastic matrices. We apply the results to minimax theory and provide asymptotically optimal strategies under $c$-interior PML constraints for binary hypothesis testing and mean estimation. The results show that disclosure control with PML allows analysts to reason about systems in a more differentiated manner: For example, it allows us to quantify the privacy leakage of deterministic systems, and can give precise adversarial guarantees with respect to arbitrary distributional assumptions. Interestingly, a recurring theme in the disclosure analyses is that if the privacy problem is relatively regular (if the density bound $c$ is large), private inference can be possible without incurring any additional cost in terms of sample complexity.
Yihan Zhang
Fri 25 Sep
math.PRcs.DScs.NAmath.NAmath.STstat.TH
pdf html
- reading eigenvalue, rich, matrices, theory
- interest Computational methods for heterogeneous agent and dynamic economic models 0.80
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.81
Abstract
We initiate the study of approximating the top eigenvalue and eigenvector of a random symmetric matrix $ A \in \mathbb{R}^{n\times n} $ using $ q(A)b $ where $q$ is a degree-$d$ polynomial and $b$ is a standard Gaussian vector independent of $A$. For spiked GOE $ Y = \lambda vv^\top + X $, we identify $ d_\star = \frac{\log(n)}{2\log(\lambda)} $ to be the critical degree threshold above which accurate approximation of the top eigenvalue and eigenvector is possible. This sharpens the common belief that spectral methods can be implemented by $ O(\log(n)) $-step power iterations and offers a precise connection between spectral methods and low-degree polynomial algorithms, a popular proxy for all polynomial-time algorithms. For GOE $X$, we identify $ d_\star = n^{1/3+o(1)} $ to be the critical degree threshold for top eigenvector approximation, whereas constant degree suffices for top eigenvalue approximation. Moreover, in the limit where $ d/n^{1/3} $ converges to a positive finite constant, we compute the exact asymptotic eigenvector approximation accuracy in terms of the expected squared overlap. These results significantly improve upon predictions made in randomized numerical linear algebra for deterministic data matrices that the iteration count of power methods is governed by the inverse spectral gap. Technically, our analyses leverage extremal properties of Chebyshev polynomials and draw upon the rich literature of random matrix theory.
Karolina Gliszczyńska-Schroeder
Tue 29 Sep
stat.MEecon.EM
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.84
- reading bayesian causal, causal, units, treatment
- interest Causal inference and identification in econometrics 0.80
Abstract
We study the impact of conditional complier average causal effect (CCACE) estimation methods on the performance of subgroup discovery and heterogeneous causal effect estimation under imperfect compliance. Building on the Bayesian Causal Forest with Instrumental Variable (BCF-IV) (Bargagli-Stoffi et al. (2022)) method, we introduce a two-step, model-agnostic approach that allows any suitable machine learning method to be used for the CCACE estimation in the first step. Specifically, we implement two non-Bayesian tree-based methods, both using forest-based learners: DRRF-IV, a debiased transformed-outcome regression-forest approach, and a GRF-based IV adaptation of the generalized random forest framework (Athey et al., 2019). Through a simulation study, we assess the precision, bias, and the ability to correctly identify the underlying subgroup structure of the proposed methods relative to BCF-IV. The results show that the non-Bayesian methods perform competitively across the considered simulation settings, with performance improving for larger sample sizes and moderately large treatment effects, while reducing computational runtime. We apply our new methods by revisiting an empirical study that examines the effect of prompt admission to intensive care units (ICU) on 28-day mortality across 48 UK National Health Service hospitals. While previous work finds no significant overall treatment effect, we investigate whether subgroups of patients may benefit more from prompt ICU admission.
Martijn Gösgens, Bart P.G. van Parys, Bert Zwart
Mon 28 Sep
math.STmath.PRstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.88
- favorites near “Optimal uniform convergence rates and asymptotic normality for series estimators under weak dependence and weak conditions” 0.86
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.81
Abstract
To design systems that are protected against events much rarer than the observational record, extreme-value methods are needed to extrapolate distribution tails. Tail-index estimators such as the Hill estimator are central to this extrapolation, but overestimating the tail exponent can lead to substantial underestimation of rare-event probabilities. Motivated by this, we derive large- and moderate-deviation asymptotics for the Hill estimator and use them to construct estimators whose probability of exceeding the true tail index decays at a controlled exponential rate (the decay rate). In the large-deviations regime, we show that a simple rescaled version of the Hill estimator achieves an optimal balance between bias and decay rate among scale-invariant estimators based on the same top $k$ order statistics. Under a second-order condition, we quantify the effect of the Hill bias, analyze a bias-corrected estimator, and identify sufficient conditions for moderate deviations in the boundary case where the second-order parameter $\rho$ equals zero.
Hamed Khosravi, Xiaoming Huo
Tue 29 Sep
stat.MEcs.AI
pdf html
- reading confidence sequences, anytime-valid, nominal, little
- interest Adaptive data collection and anytime-valid inference 0.85
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.81
Abstract
Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of models and for the set of models that may be best. For pairwise battles they also give ranks. None gives ranks when all models are scored on the same items, which makes their scores dependent. We construct rank confidence sequences: for every model, a set of ranks that contains its true rank, simultaneously for all models and at all times, at a chosen error level $\alpha$, in finite samples. The construction combines betting e-processes, one for each ordered pair of models, with closed testing over the possible orderings of the models. It allows any dependence between the models' scores on an item. The method has two advantages. A leaderboard can be inspected after every item without inflating its error rate. The evaluation of each model can stop as soon as the question asked about it is answered, which saves compute. When results are examined only once, halfway through or later, little power is lost relative to fixed-sample methods. The paper quantifies these advantages in simulations and on public leaderboard data.
Aram Ebtekar, Marcus Hutter, Danica J. Sutherland
Tue 29 Sep
stat.MLcs.AIcs.ITcs.LGmath.ITmath.STstat.TH
pdf html
- reading theory, short, infinite, inductive
- interest Online learning and statistical learning theory 0.85
- favorites near “Data Science at the Singularity” 0.84
Abstract
Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch (NFL) theorems force any learner that beats chance on some environments to underperform on others. We might hope that past experience informs which environments to expect, but NFL applies equally to meta-learning. Thus, any method that makes meaningful predictions necessarily begins with an inductive bias external to the data. Choosing to bias toward short programs yields Solomonoff induction (SI), whose performance is competitive against all computable learners - albeit up to "constants" that become large when comparing against specialized methods that exploit background information. We therefore relativize SI to an information vantage point, biasing toward short programs with access to all preexisting information. This reframes the inductive bias: instead of seeking some absolute notion of simplicity, we favor accessibility with respect to our vantage point. An algorithm can only outpredict the relativized SI to the extent that its code contains additional information about the data, and no algorithm can generate such information. While SI is incomputable and hence not a practical algorithm, it provides a formal optimum for inference in the limit of infinite compute, and there is evidence to suggest that frontier AI systems roughly approximate it. Thus, the only known answer to meta-NFL is rooted in algorithmic information theory, which we should expect to play a fundamental role in explaining the generalization behavior of modern (and future) AI systems.
Bing Cheng, Yi-Shuai Niu, Howell Tong, Shing-Tung Yau
Thu 24 Sep
stat.ME
pdf html
- reading theory, infinite, state, generic
- interest Computational methods for heterogeneous agent and dynamic economic models 0.80
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.78
Abstract
Classical continuous-time non-linear filtering fails in generic non-linear state spaces due to an infinite Lie algebraic derivative explosion ($\dim(\mathcal{E}) = \infty$), leading to filter divergence. To resolve this four-decade crisis, we introduce the SMG-Yau-Yau filter and lossless data assimilation framework built on Statistical Fiber Theory over an Orlicz manifold $\mathcal{M}$. By equipping $\mathcal{M}$ with a Riemannian submersion and an Ehresmann connection, the unconstrained Duncan-Mortensen-Zakai score velocity field is orthogonally decomposed into Statistically Verifiable Directions ($\text{SVD}\chi_f$) and Structural Internal Directions ($\text{SID}_f$). System non-linearities and unclosed Lie commutators are orthogonally quarantined in $\text{SID}_f$, protecting macroscopic base parameters from spatial derivative pollution while preserving total score variance energy. We unify the asymptotics through a Dual-Axis Collapse Mechanism: proving our model identically recovers classical Yau-Yau dynamics when $\dim(\mathcal{E}) < \infty$, while large-sample limits ($N \to \infty$) induce a thermodynamic quench that flattens infinite-dimensional geometry under $\dim(\mathcal{E}) = \infty$. Finally, we formulate the Active Acausal Tension (AAT) functional to monitor accumulated model misspecification online. Exceeding a topological capacity threshold triggers Gauge Symmetry Breaking (GSB), which dynamically expands base coordinates ($d \to d+1$) to ensure non-asymptotic stability and convergence to the exact state density.
Carlos A. Martinez, Nelson A. Cruz, Sandra E. Melo, Oscar O. Melo
Tue 29 Sep
stat.ME
pdf html
- reading empirical, linear models, observational, theory
- interest Causal inference and identification in econometrics 0.84
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.82
Abstract
The analysis of experimental or observational data often requires controlling for the effects of variables of secondary interest. Two common approaches to account for such parameters are grouping experimental or observational units into homogeneous groups (blocking), and observing quantitative variables on each one of them (covariance analysis). Often, both can be used and; therefore, comparing their performance poses a research question. Specifically, under randomized complete block designs where the blocks are defined by quantitative variables, there is interest in comparing the two-way Anova (TWA) with analysis of covariance (Ancova) models. Motivated by real-life problems coming from statistical consulting, we compared the performance of linear models considering block effects and those considering covariate effects from analytical and empirical perspectives. Models assumed Gaussian errors and included the effects of T treatments. The formal comparison relied on the theory of orthogonal projections and vector spaces under containment relationships of column spaces of the design matrices induced by conditions on the experimental design or data structure. We found asymptotic and small-sample conditions for the TWA to be at least as precise as the Ancova, as well as for these models to be equivalent. As to the empirical perspective, models were compared through a simulation study motivated by agronomical experiments. The results suggested that when the covariates show complex heterogeneity patterns, the Ancova outperforms the TWA with blocks defined by traditional approaches in terms of precision of location parameters, but not in hypothesis testing accuracy.
Bernard Colin
Fri 25 Sep
stat.ME
pdf html
- reading data analysis, functional data, functional, overview
- interest Kernels, operators and functional data 0.83
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.80
Abstract
In this paper, one provides a comprehensive mathematical and practical overview of Functional Data Analysis (FDA) specifically applied to parametrized curves in Rp. One observes that curves depending on continuously parameter are naturally present across many fields, such as, for example, monitoring child development in pediatrics, assessing meteorological phenomena, analyzing financial portfolios, tracking neurological functions, and mapping geographic pollution levels. One considers a formal theoretical framework given by a Cartesian product of real separable Hilbert spaces, to model these curves mathematically. To bridge the gap between discrete experimental measurements subject to errors, and smooth continuous functions, the paper details the essential phase of data smoothing and fitting and sets its mathematical formulation via of the Ordinary or Penalized Least Squares criteria which, for the later, using the Sobolev spaces framework, incorporates a smoothing parameter ${\lambda}$ and differential operators to prevent erratic geometric behaviors by penalizing excessive curve roughness. Furthermore, one extends the usual Principal Component Analysis to its functional counterpart in $\mathbb{R}^3$, using the calculus of variations and the Euler-Lagrange multiplier theorem in order to find eigenfunctions and eigenvalues of the covariance operator to exhibit the optimal decomposition of the spatial variance and finally, one highlights the geometric advantages of FDA over traditional multivariate data analysis, emphasizing its unique capacity to capture critical differential features as velocity, acceleration, curvature, and torsion.
Eric Frankel, Banghua Zhu, Sewoong Oh, Lillian J. Ratliff
Tue 29 Sep
cs.LGcs.AIstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.86
- reading tuned, language model, style, feedback
- interest Language models as statistical objects 0.83
Abstract
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF, DPO, and GRPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an increasing scale. Our code is available at https://github.com/SewoongLab/abc-align .
Andrea Nava, Ken Rölli, Armin Begic, Fabio Sigrist
Thu 24 Sep
stat.MLcs.LG
pdf html
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.85
- reading processes, stochastic process, irregular, discontinuities
- interest Language models as statistical objects 0.83
Abstract
Neural Processes (NPs) are model-based meta-learners that implicitly learn a stochastic process and adapt to a new task from a small context set. Most extensions of NPs focus on improving the neural network architecture. We instead develop an extension motivated by the shared hierarchical interpretation of meta-learning and mixed-effects models. Specifically, we introduce Neural Process Boosting (NPBoost), which decomposes structured response variability into tree-boosted fixed effects shared across tasks and NP random effects that capture stochastic task-to-task variation. We propose to train the two components jointly using a boosting algorithm in which an NP learns residual task-specific structure and a tree ensemble estimates common patterns across tasks. Across synthetic and real-world tabular meta-learning problems, this decomposition improves over a standard NP when the shared structure contains discontinuities or other irregular patterns that boosted trees can represent effectively.
Harold Tankpinou Zoumenou, Simon Ferreira, Charles Assaad, David Hajage, Fabrice Carrat, Alexandra Beurton, et al.
Tue 29 Sep
stat.ME
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.82
- reading treatment, rise, causal, actually
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.86
Abstract
In target trial emulation, the treatment strategies that patients follow are inferred from the treatments they actually receive in routine care. However, in most settings, the outcome may preclude the observation of planned treatment, giving rise to immortal time bias through misclassification of treatment strategy, while markers of treatment response may influence subsequent treatment decisions, giving rise to time-varying confounding. A key step toward unbiased treatment effect estimation is to partition follow-up into sufficiently short time intervals to unfold the feedback relationships involving treatment and represent the resulting causal relations with a directed acyclic graph. In this study, we present the possible within-interval causal orderings induced by this partitioning, discuss their causal implications, and assess their plausibility across clinical settings. For each causal ordering, we derive the corresponding g-formula. Using ancestral multi-world networks and simulations, we show that the standard cloning-censoring-weighting estimator is invalid when treatment affects the outcome within a time interval, and we propose a modified version of the method that restores its validity in this setting. Finally, we analyze the consequences of choosing time intervals that are either excessively wide or excessively narrow, thereby formally establishing the need for time partitioning and providing practical guidance for selecting an appropriate partition based on the clinical setting.
Gwendal Debaussart-Joniec, Théau Blanchard, Argyris Kalogeratos
Wed 30 Sep
stat.MLcs.LG
pdf html
- favorites near “Hypoelliptic diffusion maps I: tangent bundles” 0.90
- reading operator, operators, laplacian, diffusion maps
- interest Kernels, operators and functional data 0.82
Abstract
Many datasets carry an intrinsic directionality: citations point backward in time, cells differentiate along lineages, and traffic follows preferred routes. Spectral embedding methods, including most of their extensions to directed graphs, discard this information: they symmetrize the data and map it into a Euclidean space where asymmetry cannot be represented. We instead model directed data as sampled from a Finsler manifold, whose distance depends on the direction of travel, and study the kernel operator built from this asymmetric distance. Through a moment expansion of this operator, we show that its symmetric and antisymmetric parts separate geometry from direction. As the bandwidth of the kernel vanishes, the symmetric part converges to a weighted Laplacian, recovering diffusion maps in the Riemannian case, while the antisymmetric part converges to a first-order transport operator that encodes the directionality. We prove that the corresponding graph operators, built from finitely many samples, converge uniformly and almost surely to these limits. For Randers metrics, this vector field is explicit and yields an embedding algorithm recovering both the manifold structure, from the spectrum of the symmetric part, and the underlying drift. We illustrate the approach on synthetic directed graphs and point-clouds.
Sunmin Oh, Seungsu Han, Gunwoong Park
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.88
- interest Adaptive data collection and anytime-valid inference 0.83
- favorites near “Post-selection inference for causal effects after causal discovery” 0.93
Abstract
What determines the unavoidable sample cost of learning cyclic causal structure? For cyclic linear non-Gaussian models, we study exact condensation recovery from observational data: identifying the strongly connected component (SCC) partition and all edges between components. We establish the first information-theoretic lower bounds on sample complexity for this target. For $p$ variables, maximum SCC size $s_{\max}$, and maximum external-parent count $d_B$, any estimator requires order $s_{\max}\log(ep/s_{\max})+d_B\log(ep/d_B)$ samples in the worst case over a regular model class. These bounds distinguish the costs of SCC membership and external-parent selection. Under principal invertibility and without correlation faithfulness, we establish a population block-exogeneity principle that identifies unknown root SCCs through residual independence and inclusion minimality. A sparse-adjustment characterization shows that small adjustment sets suffice to identify SCCs and their direct external parents, without regressing on all previously recovered variables. These characterizations yield BlockExo, which attains a structurally matching sample bound without knowing $s_{\max}$ or $d_B$ under suitable conditions. Simulations support the structural dependence of our sample bound and demonstrate BlockExo's sample-efficient recovery in comparisons with other methods for cyclic causal discovery.
Thiago Ramos, Alek Fröhlich, Daniel Perazzo, Massimiliano Pontil
Tue 29 Sep
stat.MLcs.LG
pdf html
- reading functional, hilbert space, hilbert, operators
- interest Kernels, operators and functional data 0.84
- favorites near “Operator-valued Kernels for Learning from Functional Response Data” 0.84
Abstract
We introduce the Functional Spectral-Newton Method (FSNM) for learning the leading singular structure of a conditional expectation operator without fixing a basis or reproducing kernel Hilbert space. FSNM fits a low-rank representation of the centered joint-to-product density ratio kernel by alternating functional Newton updates. Each update reduces to a preconditioned regression, which we approximate with vector-valued regression trees in a stagewise boosting procedure. At the population level, we establish descent and an $O(1/T)$ best-iterate block-stationarity rate under a relative weak-learner accuracy condition, and show that every nondegenerate local minimum over the full centered $L^2$ spaces is a globally optimal rank-$d$ approximation. Synthetic experiments show that FSNM recovers a low-rank density ratio and its leading spectral structure, and that the same learned kernel can answer multiple conditional queries without refitting.
Zhuoer Shen
Tue 29 Sep
math.STstat.MLstat.TH
pdf html
- interest Kernels, operators and functional data 0.88
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.84
- reading theory, structural, signal, minimax rate
Abstract
Motivated by blind denoising in diffusion models, we study estimation of an unknown Gaussian noise level from a single high-dimensional observation, assuming the signal law P is known. We characterize the minimax mean-squared error under two structural assumptions on P. For signals with covering complexity k, the minimax rate is $\widetilde{\Theta}_\Lambda(\min\{\Delta_\Lambda^2,d^{-1}+k^2d^{-2}\})$, and the constrained MLE attains it up to logarithmic factors. For $\alpha$-strongly log-concave signals, the rate is $\Theta_\Lambda(\min\{\Delta_\Lambda^2,(1+\alpha^{-1})^2d^{-1}\})$, attained up to constants by a $P$-centered norm estimator. These results show that the structure of the signal law determines both the difficulty of blind noise estimation and the appropriate estimator.
Ashwin Renganathan, Peter Bachman
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.84
- favorites near “Preference-based Online Learning with Dueling Bandits: A Survey” 0.85
- interest Adaptive data collection and anytime-valid inference 0.82
Abstract
Multiobjective Bayesian optimization (MOBO) with Gaussian process (GP) surrogates is a sample efficient approach to solving multiobjective optimization problems. In MOBO, a Bayesian decision theoretic acquisition function guides the adaptive selection of new candidate inputs, on which objectives and constraints are evaluated to update the surrogate model sequentially. Existing approaches maintain independent GP models for the objectives and constraints, with new observations evaluating all objectives and constraints in a coupled fashion. However, the objectives and constraints often contain inherent correlations which, if exploited, can enable ${decoupled}$ evaluations where only a subset of them are evaluated at each round. We present a new approach that leverages a multitask GP model to jointly learn all objectives and constraints, and propose a total correlation metric that enables identifying an ${optimal}$ subset of objectives and constraints to be evaluated at every round, even under uniform evaluation costs. Theoretically, we show that our acquisition policy is asymptotically consistent despite decoupling and that our proposed decoupled subset selection rule maximizes the expected posterior entropy reduction about unevaluated tasks under mild conditions. Empirically, we show that our approach outperforms coupled and decoupled baselines in the state of the art.
Jiyue Qin, Yuyao Wang, Ronghui Xu
Tue 29 Sep
stat.ME
pdf html
- reading doubly robust, semiparametric, doubly, treatment
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
- interest Causal inference and identification in econometrics 0.79
Abstract
Semi-competing risks occur when individuals may experience a non-terminal event and a terminal event, where the terminal event censors the non-terminal event but not vice versa. In the presence of covariate-dependent censoring, augmented inverse probability of censoring weighting (AIPCW) framework has not been developed for the special setting of semi-competing risks which, unlike competing risks, have asymmetric event time structure and complex estimands. Using semiparametric theory on coarsened data, we carefully develop an AIPCW framework for semi-competing risks. When treatment effects are of interest, we further integrate this approach with augmented inverse probability of treatment weighting (AIPTW), yielding a framework for estimating causal estimands under covariate-dependent censoring. The resulting estimators are shown to be doubly robust. Using the proposed framework, we estimate treatment specific risks of: i) non-terminal event, ii) terminal event without the non-terminal event, and iii) terminal event following the non-terminal event. We evaluate the finite sample performance through simulations, and apply the method to data from the Honolulu Asia Aging Study to assess the causal effects of midlife heavy drinking on late life cognitive impairment and mortality.
Ebrahim Khaled Ebrahim
Fri 25 Sep
stat.MEstat.COstat.ML
pdf html
- favorites near “Deep Learning for Individual Heterogeneity” 0.86
- interest Computational methods for heterogeneous agent and dynamic economic models 0.80
- your work near “Automated Solution of Heterogeneous Agent Models” 0.83
Abstract
Goodness-of-fit tests for logistic regression are least reliable where they are most needed: at small samples their levels drift from the nominal one, and combining them worsens the drift. We propose a test whose statistic is a convolutional network, trained once on simulated departures, that reads misfit as a picture: a grid of standardized residuals over covariate ranks. The analyst never trains. The network ships frozen, and the p-value is the rank of the observed score within the analyst's own bootstrap, so the level is a property of the calibration rather than of what the network learned. We prove exactness under pivotality, asymptotic exactness without it, and a consistency theorem whose key condition is computable from the frozen weights in one forward pass, giving a per-alternative certificate; we also measure the test's blind cone. On a pre-declared sixty-cell grid the deployed level stays in the nominal band in fifty-eight cells, while that grid's power criterion failed; on a published benchmark it is the most stable of thirteen levels across settings and sample sizes, and the test outpowers every partition test at every sample size while ranking seventh overall. A bank-failure application and a versioned release close the paper.
Chenyu Lu, Zijun Chen, Nian Si
Tue 29 Sep
cs.LGmath.OCstat.ML
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.82
- interest Adaptive data collection and anytime-valid inference 0.82
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.83
Abstract
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback--Leibler and $f_k$-divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP to be weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for estimating the robust optimal average reward and $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for learning an $\epsilon$-optimal policy. Here, $p_{\wedge}$ is the smallest positive nominal transition probability and $u_{\delta}^{\ast}$ is a robust optimal bias function. We further provide an almost-tight explicit upper bound on $\operatorname{Span}(u_{\delta}^{\ast})$. Finally, we validate the predicted $n^{-1/2}$ convergence rate through numerical experiments.
Xiaokai Luo, Chenghao Xu, Haotian Xu, Carlos Misael Madrid Padilla, Daren Wang
Wed 30 Sep
stat.MEmath.STstat.MLstat.TH
pdf html
- interest Macroeconometrics and time series 0.86
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.80
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
Abstract
We study multiple change-point detection in multivariate time series whose distributions change in a piecewise constant manner. Distributional changes can manifest across different moment orders, from shifts in the mean and covariance to changes in higher-order moments. Higher-order moments capture increasingly rich distributional features but become difficult to estimate in high dimensions. Our tensor representation unifies moments of different orders within a common linear algebraic framework, enabling a new method to detect changes in moments of all orders up to a prescribed fixed order $p$. The resulting procedure accommodates temporal dependence and allows the dimension of the time series to grow with the sample size. Under suitable regularity conditions, the proposed procedure achieves a localization error rate that matches a newly developed minimax lower bound. We further derive limiting distributions under both nonvanishing and vanishing moment jumps and construct asymptotically valid confidence intervals in the vanishing-jump regime. Numerical experiments and real-data analyses demonstrate the method's effectiveness in detecting moment changes and a range of distributional shifts.
Nils Lid Hjort
Fri 25 Sep
stat.OT
pdf html
- interest Adaptive data collection and anytime-valid inference 0.86
- reading expected value, need, question, person
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.77
Abstract
From a deck of cards, how many cards do I need to draw, until the first ace? I identify the distribution for this waiting time $T$, and its satisfyingly nice expected value ${\rm E}\,T=(N+1)/(n+1)$, with $N$ the number of cards and $n$ the number of aces; hence $53/5=10.6$ for the standard setup. After having solved this Question One I go on to certain alternative solutions and extensions, involving e.g. Beta approximations. I also consider the distributions and means for the 2nd, the 3rd, the 4th occurrences of aces, with generalisations, where there is a Dirichlet distribution in wait for us, with further links to order statistics for the uniform. Furthermore, an apparatus is developed for obtaining estimators and full confidence distributions for applications where one knows the number $n$ of aces, but not the deck size $N$; and correspondingly for inference about the unknown population size $N$ when $n$ is known. If you have 1000 people in a room, and need to interview 11 of them until you've found the first left-handed person, how may left-handed are there in the room -- here we need both an estimate and a clear measure of uncertainty.
Xiangyu Yu, Weiyu Liu
Mon 28 Sep
stat.MEstat.APstat.ML
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.86
- reading factual, versus, large, processes
- interest Causal inference and identification in econometrics 0.80
Abstract
World models for epidemiology learn from records shaped by study designs, including assignment, sampling, measurement, and related processes. A model may therefore reconstruct observed trajectories while learning an intervention contrast that depends on how records were collected. We formalize study design as constraints on a world model latent-world interface, action mechanism, observation likelihood, and target readout. This distinguishes design-respecting models, which encode these constraints, from design-ignoring models, which fit selected records without relating assignment and observation to the intended intervention question. Using a large scale cluster-randomized test-negative trial, we hold the latent structure and fitting settings fixed and compare a design-ignoring case-count model with a design-respecting test-negative observation model. Both models achieve comparable factual reconstruction. Yet across 500 paired resampling experiments at different relative sampling intensity, median contrast changes are substantial for the design-ignoring model, versus merely marginal for the design-respecting model. The same qualitative separation holds when sampling also varies across clusters. Thus, factual reconstruction alone does not establish design alignment; testing whether fitted contrasts preserve design-implied observation-process invariances provides a sharper evaluation.
Michael Björklund
Thu 24 Sep
math.STmath.PRstat.TH
pdf html
- reading frequencies, empirical, fourier, large
- interest Adaptive data collection and anytime-valid inference 0.84
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.79
Abstract
We study inference for hyperuniformity from one large realization of a stationary point process. For point processes with an essentially free translation action of completely positive entropy, empirical Fourier statistics at Lebesgue-almost every frequency have asymptotically complex Gaussian laws, without quantitative mixing or cumulant-summability assumptions. Randomly sampled frequencies then give a tractable limiting experiment for low-frequency spectral mass. If the frequencies are uniform on the Euclidean ball $B_r$ of radius $r$, the mean of the limiting squared Fourier statistic is \[ A_r=\frac{\sigma_\eta(B_r)}{\rho_\eta\lambda_d(B_r)}. \] Hyperuniformity is characterized by $A_r\to0$ as $r\downarrow0$, so no continuity assumption on the structure factor at the origin is needed. Under \[ A_r=s+c r^\alpha+O(r^\beta),\qquad 0<\alpha<\beta, \] with a specified bound on the remainder and local square-integrability of the normalized Bartlett density, a two-radius extrapolation eliminates the $r^\alpha$ term and estimates $s$ with an explicit $O(r^\beta)$ bias bound. This yields confidence bounds and a one-sided test of $H_0:s=0$ whose asymptotic rejection probability under each fixed null process does not exceed the chosen significance level.
Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, Cheng Zeng
Mon 28 Sep
cs.LGcs.AIstat.ML
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.83
- interest Reinforcement learning, policy learning and off-policy evaluation 0.81
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.79
Abstract
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interface inherits the same lower bound. A posterior-coverage theorem characterizes budgeted retesting, while an exact edit decomposition shows that the shifted set is the unique support of an error-free target update. Causal Core implements these conditions through evidence-gated writing, readout filtering, temporal credit, hidden-context setup, and local diagnostic updates. Experiments cover finite causal systems, continuous simulators, an official TD-MPC2 world model, and Qwen2.5-7B-Instruct. A frozen Qwen last-layer probe reaches 0.958 balanced accuracy on source mechanisms but 0.583 on changed delays; the gated mechanism state reaches 1.000 and accepts only 0.056 of synchronized-readout candidates. In TD-MPC2, five target states per actuator recover effect-sign accuracy from 0.057 to 0.948 without degrading stable responses. Causal retention is therefore distinct from task sufficiency and source-domain decodability.
Jorge P. Arpasi
Mon 28 Sep
math.HOmath.MPmath.STstat.TH
pdf html
- reading technique, purpose, required, density
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.78
- interest Macroeconometrics and time series 0.81
Abstract
The purpose of this article is to explicitly describe how the integration technique known as the change of variables is used in graduate-level Probability and Statistics courses. Throughout the article, we employ this integration technique seven times across various examples and results. We demonstrate that four applications of this technique are required to derive the probability density function (PDF) of the Student's t-distribution.
Jason Wein, Stephan Wojtowytsch
Wed 30 Sep
math.NAcs.NAmath.CAstat.ML
pdf html
- reading double descent, interpolation, basis, double
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.82
- interest Kernels, operators and functional data 0.79
Abstract
The Runge phenomenon in polynomial interpolation is often considered a classical analogue of the double descent phenomenon in machine learning. In this note, we explore overparameterized polynomial interpolation in three popular polynomial bases: Monomial, Chebyshev and Legendre basis with coefficients that are minimal in the $\ell^2$-norm (and, for the monomial basis, also those minimal in the $\ell^1$-norm). We present our results primarily for equidistant and Chebyshev data points, but many results are independent of the exact form of sampling.
Fernando Pires Hartwig, George Davey Smith, Frank Dudbridge, Jack Bowden
Wed 30 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.88
- reading empirical, causal, large, heteroscedasticity
- interest Causal inference and identification in econometrics 0.80
Abstract
Mendelian randomisation (MR) implemented through instrumental variable (IV) analysis is a popular strategy for strengthening causal inference in observational studies. A key assumption for several MR estimators, including MR-Egger regression, is the INstrument Strength Independent of Direct Effect (INSIDE) assumption. However, there is no established empirical test for assessing the plausibility of this assumption. Moreover, INSIDE depends on how genetic variants are coded (i.e., on the choice of the effect allele), which is often arbitrary and therefore hampers assessing the plausibility of this assumption on substantive grounds. In this paper, we show that the all-positive coding scheme (i.e., for all variants, choosing the allele positively associated with the exposure as the effect allele), which is typically used in MR-Egger, is equivalent to a coding-invariant model that can be given a natural interpretation because the direct effect parameters under this coding scheme are in the same direction as the bias of individual-variant ratio estimators. Moreover, using both theoretical arguments and simulations, we show that, under commonly assumed data-generating models in the MR methodological literature, heteroscedasticity of instrument-outcome coefficients according to instrument-exposure coefficients is a feature of at least some types of INSIDE violation, indicating that heteroscedasticity tests could contribute to assessing the plausibility of the INSIDE assumption. We further highlight specific cases where the test would not work. We illustrate its application by re-analysing a real dataset assessing the causal effect of large particle high density lipoprotein cholesterol on age-related macular degeneration.
Sam Power
Thu 24 Sep
stat.CO
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.81
- interest Adaptive data collection and anytime-valid inference 0.83
- reading uniformly, applies, perturbation, strongly
Abstract
The multivariate expectile-based distribution is a `spiked' perturbation of a multivariate Gaussian distribution, proposed by Arbel et al. in 2023. The rejection sampler proposed for this distribution in the initial work is valid, but its acceptance probability can degenerate badly both in high dimension and in the strongly asymmetric limit. We give a simple alternative. After affine whitening and a polar decomposition, the sampling problem reduces exactly to a univariate distribution, and an additional hyperbolic change of variables exhibits this distribution as log-concave, so that a universal and uniformly efficient construction of Devroye applies. The resulting exact sampler therefore enjoys an acceptance probability of at least $1 - \exp \left( -1 \right) = 0.632120 \ldots$, uniformly over the dimension and all admissible asymmetry parameters.
Fenglin Yang, Haoran Lei, Yan Chen, Jin-Hong Du
Tue 29 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.86
- interest Semiparametric inference and debiased machine learning 0.84
- favorites near “Dynamically Optimal Treatment Allocation” 0.81
Abstract
In observational causal inference, covariate balancing is widely used to reduce source-target covariate shift, but under weak overlap in high dimensions, stronger balance can induce concentrated weights and increase variance. Balance measures how well the target covariate distribution is represented, but does not by itself determine how reliably the counterfactual mean can be estimated. We develop risk-calibrated balancing for the average treatment effect on the treated, which applies ridge augmentation to any normalised base weights and selects its penalty using conditional prediction risk of the counterfactual mean. Under a random-effects predictive model, we derive an exact finite-sample decomposition of this risk into residual covariate imbalance and weight-induced variance. For design-independent base weights under proportional asymptotics, we characterise how limiting risk depends on source and target covariance geometry, population mean shift, and weight concentration. For covariate-adaptive base weights, we develop a uniformly consistent target-aware risk estimator whose minimiser attains vanishing scaled oracle excess risk. Simulations show that the high-dimensional risk predictions remain informative for adaptive balancing and that target-aware tuning generally reduces excess target risk. Empirical analyses of job-training and single-cell perturbation data show that risk-calibrated balancing generally improves on the corresponding base estimators, with larger gains under weaker overlap.
Xinyu Song
Fri 25 Sep
econ.EM
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.76
- interest Computational methods for heterogeneous agent and dynamic economic models 0.79
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.81
Abstract
We introduce a network realized GARCH-It\^o model in which volatility transmission is a dynamic relation among the latent daily integrated volatilities of multiple assets. An unknown directed and signed network is embedded in a continuous-time variance process and appears in the resulting exponential daily recursion. Intraday returns identify integrated volatility and hence the network that governs its propagation. For a fixed network template, feasible estimation based on realized volatility is first-order equivalent to estimation based on latent integrated volatility. For an unknown network, regularization selects the relevant structure, and, under oracle conditions, a fixed-rank local refit provides conditional inference for model-implied response-based connectedness. In an application to nine U.S. sector ETFs from 2007 to 2025, the model attains the lowest average out-of-sample QLIKE among the reported recursive forecasts, although its advantage over HAR is not statistically significant. In the full sample the weighted LASSO selects an empty sparse support, so inference concerns the rank-one projected factor component; this conditional analysis identifies Energy as a net volatility transmitter. Aggregate connectedness is more stable across volatility measures than individual sparse channels.
Seiyun Shin, Juhyeong Pang, Kwang-Sung Jun
Tue 29 Sep
cs.LGstat.ML
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.83
- favorites near “Eluder Dimension and the Sample Complexity of Optimistic Exploration” 0.85
- interest Adaptive data collection and anytime-valid inference 0.83
Abstract
We study approximate sampling: given $N$ independent samples from a proposal distribution $\mu$, the goal is to select one whose distribution is close to a target $\pi$ specified only up to a normalizing constant. Block and Polyanskiy (2023) provide finite budget error bounds for approximate rejection sampling (RS) as a function of the acceptance threshold $M$. The threshold $M$ giving the smallest bound, however, depends on properties of $(\pi,\mu)$ that are typically unavailable from the observed sample. This raises a natural question: Can one attain the best RS guarantee without taking $M$ as input? We answer affirmatively by proposing a parameter-free sampling algorithm called uniform race (UR), based on importance weights, which are ratios of target to proposal probabilities (or densities). It divides each observed weight by an independent uniform random variable to form a score and returns the candidate with the largest score. For every budget $N$, its total variation error satisfies the RS upper bound for every fixed threshold $M$ simultaneously, thereby achieving the best such bound in hindsight. We also characterize its output distribution conditional on the largest score, identifying when it is exactly the target $\pi$. Uniform race has no larger total variation error than a natural budget-calibrated RS derived from Rohatgi et al. (2025) and sampling importance resampling (SIR). In particular, we exhibit instances where UR's error is exponentially smaller in $N$ than that of either baseline. Furthermore, we establish conditions under which attaining this RS guarantee for every $(\pi,\mu)$ uniquely determines the selection probabilities as those of UR. Finally, test-time scaling experiments on LLM math-reasoning tasks corroborate the theoretical comparisons and demonstrate that UR remains competitive in ground-truth accuracy without requiring threshold selection.
Jordi Valero, Josep Ginebra
Fri 25 Sep
math.STmath.PRstat.TH
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.81
- interest Adaptive data collection and anytime-valid inference 0.84
- reading stopped, stopping, finite-dimensional, long
Abstract
Statistical model transformations based on randomly stopped sums, maxima and minima are widely used to extend statistical models. We characterize the complete set of stopping models for which randomly stopped sum and extreme model transformations function as statistically stable (idempotent) model extensions. Stability requires the underlying stopping model to be closed under pgf composition. We prove that any finite-dimensional, connected stopping model closed under pgf composition is necessarily a family of random variables whose pgfs commute. Using the corresponding Koenigs function, we establish that these models form a statistical manifold admitting a global, one-dimensional parametrization $\theta = \Pr(N=1) \in (0, \theta_*]$, where the probability mass at $i$ is a polynomial in $\theta$ of degree at most $i$. Finally, we establish a duality between stopping models closed and containing the identity variable (the ones yielding stable extensions) and the set of probability distributions supported on the positive integers. These findings disprove the long standing conjecture that statistical stability occurs only under geometric stopping.
Stefano Maria Iacus, Giuseppe Porro, Haodong Qi, Devika Jain
Wed 30 Sep
stat.AP
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.82
- reading economic, aggregate, social, subjective
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.75
Abstract
Left-behind places are usually identified using economic, demographic and accessibility indicators, but these may not align with well-being and social cohesion. We combine ten years of territorial indicators with subjective measures derived from georeferenced social-media data for over 1,200 NUTS-3 regions in 28 European countries (2013-2023). Regional profiles and within-between panel models reveal uneven relationships: disadvantaged regions can display contrasting levels of well-being and cohesion, national context alters cross-country associations, and between-region differences diverge from within-region change. The findings show that aggregate measures conceal distinct forms and trajectories of territorial disadvantage.
Sami Chemlal, Thibaut Germain, Rémi Flamary, Vladimir R. Kostic, Karim Lounici
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Spectral State Space Models” 0.81
- reading function spaces, operators, operator, dynamical
- interest Language models as statistical objects 0.82
Abstract
Estimating conditional statistics and learning representations of a population of conditional distributions are central problems in many data-driven applications, including uncertainty quantification and dynamical systems analysis. Conditional mean operators (CMOs), a class of linear operators between function spaces, resolve these objectives by providing access to a broad class of conditional statistics. However, existing methods typically estimate each CMO independently or constrain it to prespecified function spaces, thereby preventing the exploitation of shared structure across related distributions. In this work, we posit that related CMOs share finite-dimensional input and output function spaces, and are specialized for each task with a linear operator mapping these spaces. Based on this hypothesis, we introduce MTL-CMO, a multi-task framework that jointly learns shared function spaces and task-specific operators across multiple datasets. We further introduce T-CMO, a transfer learning method that reuses the shared spaces to estimate, in closed form, the operator of a new conditional distribution. We establish statistical guarantees quantifying the benefits of jointly learning the shared function spaces. Our experiments demonstrate that learning shared function spaces improves uncertainty quantification across a broad range of conditional distributions and, when applied to Langevin and plasma dynamics, yields compact representations of complex dynamics that retain physically meaningful information and enable parameter identification.
Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli
Wed 30 Sep
cs.LGstat.ML
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.82
- interest Language models as statistical objects 0.84
- favorites near “A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning” 0.86
Abstract
Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a high-dimensional mixed-data regression model with a shared regression function, heterogeneous covariances and noise levels, and dataset sizes that may grow at different rates. We establish the minimax risk under an ellipsoidal parameter constraint for the general covariance structure and derive deterministic equivalents for the test error of ridge regression under commutative covariances. We then specialize to a target domain and an auxiliary domain with aligned power-law covariance spectra, where the theory yields explicit scaling laws in terms of spectral decay, target regularity, and the relative growth of the two datasets. These laws identify regimes in which combining data mixtures provably yields a faster scaling rate than using either dataset alone. In particular, improving the scaling law requires a specific interplay between spectra and relative sample sizes of the domains. Our numerical experiments on language models exhibit the same qualitative phenomenon: appropriate data mixtures yield a faster decrease in target-domain test loss than training on either domain alone.
Davide Giraudo, Sadillo Sharipov
Thu 24 Sep
math.PRmath.STstat.TH
pdf html
- reading processes, literature, invariance, motion
- favorites near “Introduction to regularity structures” 0.84
- interest Bayesian computation and structural estimation 0.81
Abstract
We establish a Donsker-Prokhorov invariance principle in some Besov spaces. Specifically, we show that polygonal line processes associated with partial sum processes of a triangular array of row-wise independent random variables converge in distribution to Brownian motion, extending earlier results in the literature.
Joshua G. Albert
Tue 29 Sep
stat.CO
pdf html
- favorites near “Geometric MCMC for Infinite-Dimensional Inverse Problems” 0.86
- reading evidence, likelihood evaluations, dimensionality, software
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.79
Abstract
Nested sampling estimates the evidence by assigning prior volumes to an ordered sequence of likelihood contours. Markov chain constrained-prior samplers, which are used in high-dimensional problems, generate many intermediate states before producing the next classic sample. These intermediate states are discarded because their correlation prevents them from being inserted into the ordered NS sequence without changing its order-statistic law. This paper introduces a novel method of using them to improve evidence estimation, by formulating NS in a Bayesian way and conditioning on phantom samples as Monte Carlo observations. We then present the open-source software package, JAXNS v3, and its implementation choices. We validate the approach on a set of problems, and identify its limitations via ablation. In our experiments, when problem structure is well resolved, conditioning on all retained phantom samples reduces log-evidence RMSE by at least $30\%$, with larger improvements at higher dimensionality. For the tested problems with unresolved structure, full phantom conditioning produces no detectable improvement or deterioration in evidence accuracy. Phantom conditioning produces overconfident evidence uncertainties. We also introduce two dynamic nested sampling allocation schemes. Evidence-improving allocation approximately halves the number of likelihood evaluations required to achieve comparable evidence accuracy relative to uniform allocation. Posterior-improving allocation doubles the classic posterior's Kish effective sample size for $17.5\%$ additional likelihood evaluations.
Shuoyang Wang, Wanyu Zhang, Lingli Yang, Kexuan Li
Wed 30 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.86
- reading treatment, randomized, nominal, measured
- interest Adaptive data collection and anytime-valid inference 0.81
Abstract
In some randomized trials the primary outcome is costly or slow to measure, while auxiliary variables that predict it are available for everyone. The outcome can then be measured in a probability sample. We study an adaptive design in which the statistician who selects outcomes to measure is blinded to treatment assignment. Outcomes from an initial random sample are used to fit pooled models for the outcome and its residual variance. These set the sampling probabilities for the remaining participants and may be refitted as outcomes accumulate. After unblinding, arm means are estimated by augmented inverse probability weighting. The estimator is unbiased for the complete-data treatment difference for any working models, and a martingale central limit theorem gives Wald intervals under repeated updating. We bound the variance lost by estimating the sampling rule. The bound is linear in the error of the fitted residual variance, quadratic when the optimal probabilities are not truncated, and relates the initial sample size to the learning rate of the models. Relative to designs using treatment assignment, the blinded design loses a term due to unequal residual variances in the two arms and a term due to the conditional treatment effect, which is second order near the null. In simulations, coverage was near nominal. Adaptive sampling was 12% to 34% more efficient than simple random sampling and needed 10% to 26% fewer measured outcomes for the same precision. In a resampling study of an antifungal trial, adaptive sampling reduced the sampling variance by 39%.
Samuel V. Singh, Mimi Zhang
Mon 28 Sep
cs.LGstat.ML
pdf html
- favorites near “Spectral State Space Models” 0.82
- interest Kernels, operators and functional data 0.83
- reading functional, functional data, functions, derivatives
Abstract
We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued functions, followed by a functional layer that projects them into a finite-dimensional Euclidean space. Unlike most neural networks for discrete-time sequences, MatFAE treats each sequence as a continuous function and can therefore encode trajectory dynamics (e.g., first-order derivatives) in its latent representations. Additionally, the morphology of the functional weights in the functional layer offers interpretability by revealing the regions of the input functional data that contribute most to the latent representations. We justify the design principles and properties of each intrinsic layer and detail how matrix factorization is handled during backpropagation. We apply MatFAE to a range of fMRI datasets, demonstrating its ability to efficiently learn informative representations from high-dimensional SPD trajectories and its practical value for real-world neuroimaging analysis.
Wei Luo, Neil K. Chada, Shijie Zhang, Lu Yu
Thu 24 Sep
math.STcs.LGmath.PRstat.MLstat.TH
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.82
- favorites near “Variational Bayes under Model Misspecification” 0.88
- interest Semiparametric inference and debiased machine learning 0.85
Abstract
Score-based diffusion models are increasingly considered in settings where the underlying data distribution may differ from the training distribution, yet existing theoretical guarantees largely focus on the no-shift setting. In this work, we study robust score estimation under Wasserstein perturbations of a reference distribution. For the Ornstein--Uhlenbeck diffusion, we show that robust estimation decomposes into two fundamental components: the statistical cost of learning the reference distribution and the intrinsic cost of distribution shift. The latter scales quadratically with the Wasserstein radius, and this dependence is minimax optimal. We construct an explicit finite-sample estimator achieving the resulting robust minimax rate without knowing the shift radius. When the reference distribution lies on an unknown low-dimensional subspace, the statistical term adapts to the intrinsic dimension while the shift cost remains unchanged. Finally, we show that the same decomposition governs positive-time reverse sampling and obtain matching minimax guarantees in KL divergence. Together, these results characterize how finite data, intrinsic dimension, and distribution shift affect the robustness of score-based diffusion models.
Sam Power
Thu 24 Sep
stat.CO
pdf html
- reading intuition, provides, uniformly, wood
- interest Bayesian computation and structural estimation 0.83
- your work near “Valid Inference with Imperfect Synthetic Data” 0.79
Abstract
The standard approach to stochastic simulation from the von Mises-Fisher distribution is a rejection sampler proposed by Ulrich and Wood. This note provides a theoretical justification of its efficiency, lower-bounding its acceptance probability uniformly over the dimension and concentration parameter. A novel interpretation of the proposal is given which provides some intuition for this efficiency.
Grigorios A. Pavliotis, Andrew M. Stuart, Andrea Zanoni
Thu 24 Sep
math.STcs.NAmath.APmath.NAstat.TH
pdf html
- reading equation, theory, nonparametric, kalman
- favorites near “Geometric MCMC for Infinite-Dimensional Inverse Problems” 0.85
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.77
Abstract
We study the nonparametric statistical inverse problem of recovering the space-dependent permittivity in a coupled Fokker-Planck-Darcy system from discrete, noisy observations of the Fokker-Planck solution. We consider the coupled parabolic-elliptic system on a bounded domain with the physically natural no-flux boundary condition for the Fokker-Planck equation and homogeneous Dirichlet boundary conditions for the Darcy equation. The inverse problem is indirect, since the unknown coefficient enters only through the elliptic equation and is observed only through its effect on the density in the parabolic equation. We first develop the analytical theory of the forward problem required for the statistical analysis, establishing well-posedness and a priori bounds uniform over the admissible set of permittivities. We then derive two stability estimates for the inverse problem, namely a Lipschitz-type forward estimate and a generalized backward estimate. Placing a rescaled Gaussian process prior on the log-shifted permittivity, we show that the posterior contracts around the truth at an explicit polynomial rate in the number of observations, with the posterior mean converging at the same rate. Numerical experiments in one and two dimensions, using preconditioned Crank-Nicolson and ensemble Kalman filter algorithms, complement our theoretical results.
Kyuri Park, Denny Borsboom, Mike Lees, Leonie Elsenburg, Gaby Lunansky, Karien Stronks, et al.
Tue 29 Sep
physics.soc-phstat.AP
pdf html
- reading economic, social, higher, interventions
- favorites near “Cause, Effect, and the Structure of the Social World” 0.81
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.80
Abstract
Groups differing in social and economic circumstances often differ markedly in depressive symptom levels, yet their estimated symptom networks show few clear differences. Such null results are often read as evidence that the groups' symptom systems are the same, but they only show that no difference was detected in how symptoms interact. How readily symptoms become present, and how this depends on the contextual conditions under which they are observed, is rarely compared. Building on a slow-fast perspective, we treat depressive symptoms as a relatively fast-changing system embedded in more persistent social, economic, psychosocial, health, and lifestyle conditions that shape symptom activation. Using a cohort of 23,689 adults from Amsterdam, we illustrate this with PHQ-9 depressive symptoms and a composite index of persistent socioeconomic, psychosocial, and health-related conditions, which we call Slow Risk Load (SRL). Participants with high SRL had substantially higher symptom levels than those with low SRL. A standard Network Comparison Test detected no overall difference in network structure, although global strength was modestly higher in the high-SRL group. Ising models, which separate how symptoms co-occur from how readily each becomes present, indicate higher activation under high SRL for eight of nine symptoms, and model comparison consistently favored group-specific activation over group-specific interactions. The slow-fast perspective reframes group network comparisons as a multiple-timescale problem in which network structure, symptom activation, and context should be considered jointly. It places symptom-focused and context-focused interventions at different layers of the same coupled system rather than in competition.
Matthew F. Dixon, Bertrand Nortier, Miquel Noguer i Alonso
Fri 25 Sep
cs.CLcs.IRstat.ML
pdf html
- reading apparently, structural, weaknesses, evidence
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.80
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.78
Abstract
AI-workflow governance cannot be reduced to checking the final answer: an apparently safe answer may rest on a fragile evidence path that ordinary evaluation cannot see, localize or govern. We call this hidden fragility a "structural iceberg": hallucinations and unsupported claims may form its visible tip, while consequential weakness remains submerged. Stochastic semantic evidence graphs (SSEGs) expose these icebergs by preserving workflow channels, propagating local uncertainty and identifying the hidden paths on which an apparently safe output depends. ALCE and RAGTruth show that visible failures at the tip -unsupported citations and hallucinated spans -rest on distinct submerged weaknesses and therefore require different interventions. Across retrieval, tool-use and controlled stress tests, SSEG localizes those weaknesses, supports targeted repair, produces no false automatic passes in 35,000 known-truth cases and reduces ToolSandbox review by 28.8% across 96 executions from two agent models. The same structural view carries into end-to-end governance: in a separately sealed 1,200-case FinGovBench study, adding SSEG to GPT-OSS-20B reduces unsafe releases from 452/660 to 8/660 while releasing all 540 safe cases and correctly distinguishing 592/600 matched workflow pairs. An unchanged-gate transfer to Qwen3-8B releases all 540 safe cases and none of 660 unsafe cases, whereas flat-UQ releases 520 unsafe cases. SSEG therefore moves governance below surface-level output checking, turning hidden evidence dependencies into auditable, path-specific decisions about intervention, revalidation and release.
Haoran Yan, Zhongjie Shi, Yuanzhe Xi, Peng Chen, Wenjing Liao
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.88
- interest Kernels, operators and functional data 0.83
- reading operator, operators, theory, functions
Abstract
Transformers have emerged as powerful architectures for learning solution operators of physical systems. Empirically the prediction error has been observed to decrease when the data size and model size increase, suggesting neural scaling behavior. Yet a theoretical understanding of such scaling laws for transformer-based operator learning remains limited. In this work, we develop a theoretical framework for characterizing the approximation and generalization errors of transformer-based operator learning. Our analysis builds on a local-to-global approximation principle that is naturally aligned with the softmax attention mechanism and yields discretization-invariant output functions. On approximation theory, we derive a universal approximation error of transformer-based operator learning for H\"older-regular operators. On generalization theory, we establish a power scaling law between the prediction error and the training data size. The rate of convergence represented by the scaling exponent explicitly reflects the dimensions of the input and output domains, the regularity of the underlying functions and operators, and crucially, the intrinsic dimension of the input function class. By exploiting this intrinsic low-dimensional structure, our analysis yields a power-law generalization rate for operator learning, in contrast to the logarithmic-type power-law rates appearing in existing analyses of operator learning with feedforward neural networks. Numerical experiments validate the predicted power-law scaling and confirm that the convergence rate varies systematically with the intrinsic dimension of the input function class.
Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang
Tue 29 Sep
cs.LGcs.CLstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.87
- interest Language models as statistical objects 0.86
- your work near “Valid Inference with Imperfect Synthetic Data” 0.77
Abstract
We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap uses a surrogate LLM and constructs a geometry-aware mapping from a lower-dimensional latent reference space into its token-embedding space to induce an explicit and reproducible distribution over input prompts. A response-level performance function is utilized on the response to quantify behavior severity. This enables sequential rare event simulation that concentrates evaluations on progressively more severe behaviors while preserving probability under the induced prompt distribution, which would otherwise be prohibitive to measure. Across 10 open-weight and two frontier models (GPT-5.4 and Claude Sonnet 4.6), we find that RareTrap successfully induces severe resource consumption behaviors and computes their probability with as few as 200 evaluations. RareTrap provides model developers a principled approach for evaluating language models under a common distribution, and prioritizing alignment effort to improve safety and mitigate risks.
Louis Grenioux, RuiKang OuYang, Luhuan Wu
Tue 29 Sep
stat.MLcs.LG
pdf html
- reading equilibrium, mechanics, state, consistent
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.79
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.80
Abstract
Generating independent, equilibrium samples of molecular systems at scale remains a central obstacle in computational statistical mechanics. Boltzmann Generators address this by pairing a generative model with importance sampling to obtain consistent samples from the target distribution. We introduce Normalizing Flow Flow Maps (NF$^2$M), which combines the strengths of recent stochastic flow maps with the tractability of classic normalizing flows to build a Boltzmann Generator. Unlike most methods, which correct the generative model only at the end, NF$^2$M reweighs each denoising transition as generation proceeds, avoiding wasted compute on trajectories that are ultimately discarded. At each denoising step, a conditional normalizing flow proposes clean configurations given the current noisy state (a simpler task than sampling directly from the target) and its exact likelihood enables correcting each proposal toward the true denoising transition of the target Boltzmann distribution. This is in contrast to most existing methods, whose likelihoods are approximate or expensive to evaluate, undermining the statistical reliability of the correction. We establish consistency of the corrected transitions and bound how approximation errors propagate through the sampling chain. We evaluate NF$^2$M on peptide systems, demonstrating improved sampling efficiency and sample quality.
Durba Bhattacharya, Sourabh Bhattacharya
Tue 29 Sep
stat.MEmath.STstat.APstat.TH
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- interest Semiparametric inference and debiased machine learning 0.83
- reading dirichlet process, nonparametric, slower, overcomplete
Abstract
We propose a Bayesian nonparametric factor model that infers the number of factors, induces row-wise sparsity, and merges redundant dictionary elements via exact clustering. A Dirichlet process prior is placed on the columns of an overcomplete loading matrix and fully marginalized to an exact P\'olya urn, avoiding stick-breaking and auxiliary variables. A spike-and-slab base measure allows entire factors to be exactly zero. The model uniquely combines exact zeros, exchangeability over columns, and exact merging within a single marginalized Dirichlet process, unlike CUSP, MGP, or the beta process. An exact Gibbs sampler with canonical relabeling and parallel C/MPI implementation is developed. We prove posterior contraction at rate $\sqrt{M s_0 \log n / n}$ for the covariance matrix, and in the fixed-dictionary setting obtain the minimax optimal rate $\sqrt{s_0 \log n / n}$ plus underfitting consistency; the overfitting direction is an open conjecture. The spike-and-slab is essential: without it the effective dimension scales as $pM$, yielding a slower rate. Simulations show the method is the only fully adaptive approach to recover the true rank, achieving the smallest covariance, loading, and signal-reconstruction errors, beating an oracle baseline. On van 't Veer breast cancer data ($n=97$, $p=1213$), the posterior concentrates on eight interpretable programmes; seven pass coherence and two pass Bonferroni-corrected Hallmark enrichment.
Vincent D. Zaballa, Elliot E. Hui
Wed 30 Sep
cs.LGq-bio.QMstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- interest Adaptive data collection and anytime-valid inference 0.84
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.84
Abstract
Simulation-based inference is challenging when many heterogeneous observations must be composed, hierarchical latent structure must be preserved, and the simulator is misspecified relative to observed data. We develop sampling and fine-tuning methods for diffusion-based inference in design-conditional settings, where the same simulator is queried across different experimental conditions $\xi$. We extend compositional score-based inference with a continuous-time diffusion coefficient that accounts for the number of observations, avoiding Jacobian and auxiliary-covariance corrections. We introduce Hierarchical Blockwise Diffusion Sampling (HBDS), which infers shared parameters and group-specific latent states using a single pretrained model, with the hierarchy specified only at sampling time. Together, these methods support variable observation sets and groupings without retraining. To address misspecification, we introduce path-regularized fine-tuning that adapts the learned likelihood to observations and transfers corrections to posterior inference. Using Girsanov's theorem, we quantify path divergence between pretrained and fine-tuned models across experimental designs and interpret it alongside predictive errors to distinguish candidate misspecification correction from unnecessary adaptation. We evaluate compositional sampling on exact-score Gaussian and Simple Likelihood, Complex Posterior benchmarks, HBDS with analytic and learned scores on a controlled hierarchical model, and fine-tuning and localization on a separate analytic model with known design-dependent discrepancy. Finally, we apply the framework to 940 measurements across four cell lines in a mechanistic Bone Morphogenetic Protein signaling model, where fine-tuning improves posterior-predictive accuracy relative to the pretrained model and shifts posterior marginals toward the least-squares reference while retaining spread.
Michael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry, Aryo Lotfi, Abbas Kazerouni, et al.
Wed 30 Sep
cs.LGcs.AIcs.CLstat.ML
pdf html
- reading policy, attempts, thinking, failed
- your work near “Timing as an Action: Learning When to Observe and Act” 0.80
- favorites near “Reinforcement Learning: An Overview” 0.85
Abstract
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.
Preben Forer, Urte Adomaityte, Pierpaolo Vivo
Thu 24 Sep
cond-mat.dis-nncond-mat.stat-mechmath.MPstat.ML
pdf html
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.84
- reading eigenvalue, equations, rank-one, signal
- interest Kernels, operators and functional data 0.80
Abstract
We study the high-dimensional recovery of a signal vector $\mathbf{x}$ in the presence of sparse Wishart-like noise. We define an $N \times N$ matrix $A = J+(\theta/N)\mathbf{xx}^{\top}$, where $\mathbf{xx}^{\top}$ is the rank-one deformation of the random noise matrix $J$. We consider a Wishart-like matrix $J={X}^{\top} X$, where $X$ is a sparse $M \times N$ random matrix with entries $X_{ij} = c_{ij}W_{ij}$, with $c_{ij}$ regulating the density of non-zero elements, and $W_{ij}$ the bond weights. Using the replica method, we compute analytically the top eigenpair statistics of $A$, and their dependence on the signal strength $\theta$, the rectangularity ratio $\alpha=\sqrt{M/N}$, and the average connectivity of the noise. The spectral observables are expressed in terms of a system of Recursive Distributional Equations, which are efficiently solved via a Population Dynamics algorithm. They allow us to compute the average largest eigenvalue $\langle\lambda_1\rangle_{A}$, the average top eigenvector component density, and the average overlap between the top eigenvector of $A$ and $\mathbf{x}$. We identify a critical threshold $\theta_{\mathrm{crit}}$--depending on the average connectivity of the noise--that marks a BBP-like phase transition: below this value, $\langle\lambda_1\rangle_{A}$ is unaffected by the signal, and the overlap vanishes. Thus, the signal is not recoverable from the top eigenvector of $A$. For $\theta>\theta_{\mathrm{crit}}$, the signal-related outlier eigenvalue becomes $\langle\lambda_1\rangle_{A}$ and the overlap is nonzero, allowing for recovery of the signal. The results are in excellent agreement with numerical diagonalisation. We show that in the dense limit, the recovery threshold and eigen-statistics converge to the results predicted by the classical BBP transition for additive rank-one deformations of dense Wishart matrices.
Attila Lovas, Lóránt Nagy
Wed 30 Sep
stat.MLcs.LG
pdf html
- reading empirical, equation, computing, fails
- interest Computational methods for heterogeneous agent and dynamic economic models 0.78
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.79
Abstract
Bre\v{s}ar and Mijatovi\'c \cite{bresar2025} show that Ornstein--Uhlenbeck diffusion is hard to beat in forward convergence under assumptions that exclude superlinear drift. We instead test superlinear Langevin diffusions for score-based image generation, computing their conditional scores numerically from a Fokker--Planck equation. In our experiments, the superlinear models beat the Ornstein--Uhlenbeck baseline on empirical Wasserstein distance across nearly the entire tested grid and show less variation across diffusion horizons. The ``hard to beat'' verdict of \cite{bresar2025} thus fails to be universal.
Ruishuo Chen, Weijia Li, Xun Wang, Yu Chen, Leheng Cai, Longbo Huang
Tue 29 Sep
cs.LGstat.CO
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.82
- reading matrices, sequential, social, policy
- interest Language models as statistical objects 0.80
Abstract
In ecology, psychometrics, and the analysis of social and financial networks, binary matrices are often analyzed conditional on their observed row and column sums, which restricts the problem to a finite sample space of matrices with the same margins. Two fundamental problems are to count this space and to sample uniformly from it. Sequential importance sampling (SIS) addresses both with independent weighted samples and an unbiased count estimator, but its efficiency depends critically on the proposal distribution. Existing proposals are analytically designed, and their accuracy can vary substantially with the margins. We show that the ideal SIS proposal, under which every weight equals the count and the variance vanishes, is exactly the policy of a generative flow network (GFlowNet) with unit reward on every matrix that has the given margins. We therefore propose MarginFlow, a framework that turns the design of the proposal into a learning problem and amortizes it across margins by exploiting their self-similarity. Every partial matrix is itself an instance with reduced margins, so one set transformer that reads the remaining margins serves every margin. We train MarginFlow on a pool of 1904 margins and evaluate it zero-shot on 1190 held-out margins, synthetic and real, from $3\times3$ to $870\times6$. On 1187 of the 1190 margins it matches or beats the best of 31 analytically designed configurations, chosen post hoc for each margin, and its median effective sample fraction is 99.8%. On the 56 margins where that best loses more than one nat of effective sample size, MarginFlow wins every one and raises the median effective sample fraction from 10.3% to 94.1%.
Henry Pritchard, Rahul Parhi
Tue 29 Sep
math.OCeess.SPstat.ML
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.84
- reading operator, operators, functions, variational
- interest Kernels, operators and functional data 0.82
Abstract
Plug-and-Play (PnP) methods replace proximal operators with learned denoisers, which produce state-of-the-art reconstruction quality, but sacrifice the variational interpretation and convergence guarantees of proximal methods. Learned Proximal Networks (LPNs) address this problem by designing the denoiser architecture so that it is exactly the proximal operator of a regularizer. In this work, we clarify the theoretical foundations of LPNs and extend the framework to a broader class of activation functions. We then study two practical mechanisms for controlling the learned prior: (i) averaging the LPN with the identity, a common heuristic in PnP methods, and (ii) directly scaling the implicit regularizer induced by the proximal operator. In particular, we characterize the regularizer induced by averaging and develop a convergent method for evaluating the scaled proximal operator.
Zhongxuan Liu, Yue Kang, Thomas C. M. Lee
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Preference-based Online Learning with Dueling Bandits: A Survey” 0.87
- interest Reinforcement learning, policy learning and off-policy evaluation 0.83
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.83
Abstract
Low-rank matrix and tensor bandits exploit structured interactions but typically assume a known reward link. Recent single-index bandit methods accommodate unknown links without directly exploiting matrix or tensor rank. We address this gap by studying stochastic matrix and tensor bandits with an unknown shared Lipschitz link and a low-rank index parameter under known regular candidate distributions and finite-variance noise. For monotone links, T-ESTOR combines robust, rank-adaptive Stein estimation with epoch-based greedy selection. Under exact selected-score access and a uniformly positive selected-design Stein signal, it achieves square-root regret with dimension dependence determined by the low-rank structure. For every admissible design, the monotone lower bound matches the rank, dimension, and horizon dependence up to logarithmic factors at large horizons, for fixed menu size and model/design constants. For nonmonotone links under a nonzero base-law Stein signal, T-BSTOR combines structured estimation with robust bin-based learning and attains the optimal $\widetilde{O}(T^{2/3})$ horizon rate for fixed dimensions, menu size, and model/design constants. Synthetic and CCLE-based experiments illustrate the benefits of structured estimation relative to vectorized and competing single-index baseline methods.
Isqeel Ogunsola, Olatunji Johnson
Fri 25 Sep
stat.ME
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.76
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.83
- interest Causal inference and identification in econometrics 0.80
Abstract
Estimating causal effects is challenging with observational data due to the lack of exchangeability, and is further complicated with spatial data. In real-life applications such as agricultural and environmental studies, unobserved spatial factors (spatial confounding, SC) and interactions of units in nearby locations (spatial interference, SI) commonly occur jointly, yet existing parametric methods address them separately. We develop a unified framework that treats both simultaneously. We first define the estimand precisely and give an identification result. We then study two estimators. The matching estimator iDAPS integrates neighbourhood exposure and spatial proximity into propensity score matching with balance-optimised weights. We prove a variogram bias bound showing its confounding bias is controlled by the spatial variability of the confounder at the matched distance; this justifies the composite metric, yields a consistency result, and provides a computable diagnostic. The doubly robust estimator recoverU+ augments the propensity and outcome models with the recovered confounder and the exposure; we prove it is doubly robust up to an explicit residual bias from the unrecoverable confounder component. We show that independence-based standard errors are anticonservative under spatial dependence; we provide spatial heteroskedasticity and autocorrelation consistent (HAC), block-bootstrap, and randomisation-based alternatives whose calibration we verify. Every theoretical result is validated numerically. Estimating the effect of selective catalytic and non-catalytic reduction (SCR/SNCR) technologies on ambient ozone reveals no evidence of ozone reduction, while a naive analysis would have delivered a confidently wrong sign. All methods are implemented in the developed open-source R package spaci.
Ivy Zhang, Dominik Rothenhaeusler
Mon 28 Sep
stat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.84
- interest Causal inference and identification in econometrics 0.80
- reading empirical, consistent, partially, generally improves
Abstract
Prediction models can perform poorly when the deployment population differs from the training population. Data from the target population would help, but individual-level target data may be inaccessible because of access restrictions or reporting conventions. We consider a multi-resolution setting in which individual-level data are available from a source population, while the target population is observed only through subgroup summaries. We propose SAGE, a one-step estimator that updates a source-trained predictor using a gradient estimated from these summaries. Motivated by diagnostics consistent with the random distribution shift model, we choose SAGE's step-size to account for both sampling and distributional uncertainty. Under this model, SAGE reduces mean asymptotic target excess risk relative to the source-trained predictor. We also show that, under the model, entropy-balancing weighted empirical risk minimization (EB), which reweights source observations to match the target summaries, is asymptotically equivalent to a full-step SAGE update. SAGE with the optimal step-size has asymptotic mean squared error no larger than that of EB. Across real-world datasets, SAGE generally improves on the source-trained predictor and one-step updates that ignore distribution shift, including when the random shift model only partially captures the observed shifts. Compared to EB, SAGE improves prediction more consistently across the sample size and shift settings studied.
Chen Yang, Jun Chen
Thu 24 Sep
stat.MLcs.LG
pdf html
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.85
- interest Adaptive data collection and anytime-valid inference 0.85
- favorites near “Measuring Price Selection in Microdata: It's Not There” 0.79
Abstract
LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public. We ask how many hidden variants a published margin can support while retaining statistical evidence of a provider's advantage over a fixed comparator. For a fixed candidate family under a Gaussian margin model, we derive a sensitivity curve that reports this maximum count as a function of a lower bound on within-family correlation. The relevant correlation must match the score used for ranking and the sampling model: in a controlled family, pooled item correlation is 0.90, whereas composite-score correlation is 0.46 under item resampling and 0.92 when MMLU subjects are resampled. An item-based audit of 394 adjacent-rank claims on the Open LLM Leaderboard finds that 391 lack statistical support even before accounting for selection. Among claims that pass the uncorrected test, certification can depend on assumptions about the hidden family's correlation. The resulting curves make these assumptions explicit without estimating the unobserved search size.
Minglei Yang, Sicheng He
Thu 24 Sep
stat.MLmath.MP
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.80
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
- favorites near “Spectral State Space Models” 0.77
Abstract
Generalized Langevin equations describe non-Markovian dynamics in which the evolution of resolved variables depends on their past. We propose a memory-conditioned diffusion method for learning stochastic flow maps of these dynamics from observed trajectories, without identifying a memory kernel or reconstructing unresolved variables. A compact, recursively updated bank of exponential filters enables the flow map to retain predictive history over multiple time scales without conditioning on long observation windows. The next-step distribution is conditioned on the current observation and this memory state, whose storage and update costs are independent of the history length for a fixed bank size. Predictive criteria guide the memory budget, with reference-assisted selection in the vector benchmark, and an optional linear projection further reduces the conditioning dimension. A kernel-based score estimator generates conditional samples without training a score network, and these samples are used to train a neural flow map for autoregressive simulation. Three numerical examples assess long-memory retention at small conditioning dimension, predictive compression in coupled vector dynamics, and non-Gaussian conditional distributions and intermittent events. The non-Gaussian example reproduces conditional asymmetry and burst statistics in a stochastic model of the plasma scrape-off layer.
Ralf Herbrich, Rainer Schlosser, Jan Lemcke, Johann Ukrow, Anna Kazachkova, Nicolas Alder, et al.
Fri 25 Sep
cs.LGcs.AIstat.ML
pdf html
- your work near “Automated Solution of Heterogeneous Agent Models” 0.82
- interest Computational methods for heterogeneous agent and dynamic economic models 0.78
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.84
Abstract
Approximate message passing on factor graphs underlies two dominant families of probabilistic inference algorithms: expectation propagation (EP) and variational message passing (VMP). Both methods approximate the marginal at each factor edge, forcing an iterative round-robin schedule, risking negative-precision messages, and, for VMP, collapsing to point estimates at Dirac-delta factors. We introduce Direct Message Approximation (DMA), which approximates factor-to-variable messages directly rather than the marginal. For normalisable factors, we define a consistency condition (requiring exactness when all other incoming messages are Dirac deltas) to guide message construction. We prove a master theorem (proper messages, any graph) bounding marginal KL from message KL, with three structural corollaries: Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages. Further, we prove a complementary $O(1/r^2)$ guarantee for the inherently improper backward message of the product factor, whose closed-form treatment has resisted prior work. As a concrete instantiation, we derive explicit DMA messages for the product and leaky-ReLU factors and assemble a Bayesian neural network (BNN) inference algorithm with one forward/backward sweep per training example and no gradient learning-rate hyperparameter, validating that the structural guarantees translate to predictive uncertainty that widens in data-sparse regions, including under model mismatch.
Ruben Karapetyan, Shenyuan Ma, Ales Wodecki, Jakub Marecek
Tue 29 Sep
math.STcs.LGmath.OCstat.TH
pdf html
- reading moments, moment, regular, connection
- interest Kernels, operators and functional data 0.82
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.83
Abstract
Recovering probability measures from their moments has numerous applications, esp. in connection with the method of moments in statistics and optimization. In the setting where measure need not be finitely atomic, but its support is known to be compact and semialgebraic with codimension at least one, the problem is still open. We combine moment-matrix kernel information with the Christoffel--Darboux kernel to provide a discrete approximation of the support. To validate the proposed approach, we test our algorithm on analytically computed moments and pseudo-moments arising from polynomial optimization problems without unique global minimizers. This complements well-known recent work on recovery of measures with algebraic support, where the kernel of a moment matrix can reveal polynomials vanishing on the support, and on recovery of sufficiently regular full-dimensional supports, where estimators constructed by thresholding the Christoffel--Darboux kernel are known to converge asymptotically to the support.
Diego Serrano, Eduardo García-Portugués, Ingrid Van Keilegom
Wed 30 Sep
math.STstat.MEstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.87
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.85
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.74
Abstract
We propose a general goodness-of-fit framework for distributions on separable metric spaces. Under suitable identifiability conditions, probability distributions are characterized by distance profiles, which motivates their use in goodness-of-fit testing, for simple and composite null hypotheses. For composite null hypotheses, parameter estimation is incorporated via a Bahadur-type expansion, and the asymptotic distribution of the empirical process for distance profiles is obtained under the null. We define test statistics based on this empirical process and derive their asymptotic null distributions. We further study the behavior of the proposed tests under fixed and local alternatives, establishing consistency results. Multiplier bootstrap procedures are developed, and their conditional asymptotic validity is established under both simple and composite null hypotheses. The methodology is illustrated with simulation studies and real-data applications for data on the sphere, hyperboloid, and simplex.
Bart P.G. van Parys
Wed 30 Sep
math.OCstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.81
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.83
- your work near “Automated Solution of Heterogeneous Agent Models” 0.80
Abstract
Fix a target accuracy $\varepsilon$, a gradient-noise level $s$, and a horizon $N$. We wish to design algorithms which minimize the probability of observing a suboptimality gap which exceeds the target accuracy, i.e., $\mathcal{E}_N = -\log \sup_{f,P} \mathbb{P}_P(f(x_A) - f_\star \ge \varepsilon)$, with the noise law $P$ known only to be sub-Gaussian. We single out a uniformly averaged schedule which is harmonic (of the form $h_k = R^2/(\varepsilon (N+m-k))$) and prove, via an optimized exponential supermartingale argument, that it attains the optimal exponent $\mathcal{E}_N^\star = \varepsilon^2 N (1+o(1))/(2R^2 s^2)$. Optimality is certified by a matching impossibility result: under Gaussian noise, a gradient-masking change of measure caps the exponent of every algorithm at the same leading order. In a small-noise limit, our setting degenerates into the adversarial-error model of G\"osgens and van Parys (2025) for subgradient methods.
Takashi Takenouchi
Mon 28 Sep
stat.MLcs.LG
pdf html
- reading empirical, bregman, technique, favorable
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.78
- favorites near “Variational Bayes under Model Misspecification” 0.86
Abstract
Estimation of parameter of probabilistic models is an important task in the field of machine learning.For models of discrete variables, calculation of the normalization constant of model is sometimes difficult and a lot of researches have been done to avoid the calculation of the normalization constant. In this paper, we tackle with the difficulty by combining a technique of empirical localization and a deformed Bregman divergence.The technique of empirical localization makes it possible to drastically reduce computational cost of the calculation of the normalization constant, and in addition, appropriate choice of the deformation for the Bregman divergence can invest the proposed estimator with various kinds of favorable statistical properties, such as efficiency or robustness against outlier noise.
Ansh Nagda, Alexander S. Wein
Wed 30 Sep
math.STcs.CCstat.TH
pdf html
- reading eigenvalue, possible, tradeoffs, large
- favorites near “Optimal Decision Rules for Weak GMM” 0.86
- interest Online learning and statistical learning theory 0.84
Abstract
We study the computational complexity of hypothesis testing in the spiked Wigner model, a prototypical model for detecting low-rank structure in a large random matrix. Below the "BBP" eigenvalue transition, it is expected that strong detection --- with both type I and II errors vanishing --- requires exponential time. Assuming this as a conjecture, we determine the limits of polynomial-time weak detection, exactly characterizing the possible tradeoffs between type I and II errors. Specifically, the optimal tradeoff is achieved by a particular linear spectral statistic. Thus, the question of weak detection is entirely reduced to that of strong detection. The proof builds on ideas of Nagda-Raghavendra (2025) and Moitra-Wein (2025). The low-degree likelihood ratio (LDLR) plays a key role: any test that slightly beats the LDLR can be boosted to have an even higher success probability. This leads us to establish a computational analogue of the Neyman-Pearson lemma for a subclass of additive Gaussian models: for a given super-polynomial runtime, the best possible tradeoff between type I and II errors is either the one achieved by thresholding the LDLR, or the trivial tradeoff that results from strong detection.
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, et al.
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.82
- reading diffusion models, state, ordinary differential, equation
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.77
Abstract
Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information collapse). We propose Simplex Diffusion Models (SDMs), a framework that lifts the diffusion process to the probability simplex to represent beliefs over categories. SDMs admit probability paths with closed-form reverse transitions and can be trained with a simple cross-entropy loss. Contrary to earlier proposals such as Dirichlet Flow Matching which requires integrating an ordinary differential equation, we introduce a DDIM-like sampler with a tunable level of stochasticity. Because SDMs operate on samples on the simplex, they can carry uncertainty across denoising steps, which mitigates information collapse. On OpenWebText, SDMs are competitive with strong Discrete Diffusion baselines, achieving $17.0$ GenPPL at $5.46$ unigram entropy in 64 sampling steps, close to real validation data. Even without Self-Conditioning (SC), SDMs outperform masked and uniform diffusion (with SC or predictor-corrector sampling) on code generation (TinyGSM, $T=0.1$; $49.0\%$ vs. $45.8\%$). Distilled down to 8 steps, SDMs solve $32.1\%$ of GSM8K problems, more than distilled Discrete Diffusion models with 128 steps ($21.4\%$).
Guillaume Braun, Ichiro Hashimoto, Masaaki Imaizumi
Tue 29 Sep
stat.MLcs.LG
pdf html
- interest Kernels, operators and functional data 0.87
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.82
- reading generalize, generalization, worse, implicit
Abstract
We study the generalization of spectral gradient descent (SpecGD) in overparameterized matrix classification with corrupted labels. Each input combines a shared low-rank signal with a rank-one sample-specific perturbation, referred to as a shortcut, that enables memorization but does not generalize. We contrast collapsed shortcuts, which share a singular direction, with dispersed shortcuts, which occupy distinct singular directions. Changing only this geometry can reverse the relative generalization of GD and SpecGD: collapsed shortcuts can favor SpecGD, while dispersed shortcuts can favor GD. In the dispersed regime, exact shortcut orthogonality eliminates the signal from the late-stage SpecGD direction, while vanishing random correlations collectively generate a small but generalization-relevant signal through a second-order effect. To identify the direction selected by SpecGD, which the spectral max-margin problem alone does not determine, we combine a refined analysis of its dual with the exponentiated-gradient dynamics of normalized loss weights. Finally, we show that a single SpecGD step can already interpolate and generalize well, while continued training converges to a direction with substantially worse generalization.
Madeline Navarro, Gonzalo Mateos, Samuel Rey
Tue 29 Sep
eess.SPstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- interest Computational methods for heterogeneous agent and dynamic economic models 0.78
- reading art, structural, equation, observational
Abstract
We study the problem of learning directed acyclic graphs (DAGs) with nonnegative edge weights from observational data. We propose the Nonnegative and Concomitant (NoCo) DAG estimator, which jointly recovers the weighted graph structure and the exogenous noise variances in the linear structural equation model for the observations. Different from prior art, this noise-adaptive formulation blends a smoothed concomitant lasso criterion with a simpler log-determinant acyclicity constraint that exploits nonnegativity and yields a more benign optimization landscape. Specifically, nonnegative weights allow us to impose acyclicity directly on the adjacency matrix without elementwise squaring of its entries, thus avoiding the well-documented degeneracy of the Karush-Kuhn-Tucker conditions. Computationally, we develop a method of multipliers' algorithm that accommodates both homoscedastic and heteroscedastic noise profiles. Within each iteration, we use block successive convex approximation to minimize the augmented Lagrangian, alternating between proximal gradient steps for the adjacency matrix and closed-form updates for the noise scales. Simulated experiments demonstrate NoCo's improved structural and edge-weight recovery relative to competing methods in a variety of settings, highlighting the benefits of exploiting nonnegativity along with noise adaptivity.
Martin Hanik, Christoph von Tycowicz
Mon 28 Sep
math.STmath.DGstat.TH
pdf html
- interest Kernels, operators and functional data 0.84
- reading existence, questions, estimator known, main
- favorites near “Why Does Deep Learning Work? - A Perspective From Group Theory” 0.80
Abstract
Bi-invariant geodesic regression generalizes linear regression to Lie groups. Its main feature is that it respects the symmetries of the group so that the resulting estimator is independent of arbitrary choices such as a reference frame. However, the local existence and uniqueness of the underlying estimator have not been shown until now. Furthermore, the convergence properties of the proposed algorithm for computing the estimator are not known. In this work, we investigate these questions. We prove that, locally, a unique estimator exists and give explicit bounds on the size of this neighborhood. We also show that the proposed iterative algorithm converges linearly to this estimator.
Stylianos Kampakis
Wed 30 Sep
stat.OT
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.79
- reading economic, gains, adjustment, empirical
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.72
Abstract
We evaluate the road safety implications of Low Traffic Neighbourhoods (LTNs) and blanket 20 mph speed limits in London using injury-collision records (2020-2024). Because London implemented 20 mph limits primarily via sign-only orders without physical traffic calming, this work evaluates the blanket statutory model rather than lower urban speeds in principle. For LTNs, we apply a causal pipeline combining naive pre/post comparisons, Empirical Bayes shrinkage for Regression to the Mean (RTM), and spatial Difference-in-Differences. Naive comparisons suggest substantial crash reductions, but Empirical Bayes adjustment reveals much of this reflects mean reversion. After correction, only one zone exhibits a statistically significant inside-area reduction. We find no consistent evidence of boundary-road crash displacement nor uniform inside safety gains. Given the modest corrected benefits, small journey-time increases on boundary roads (approximately 10 seconds per vehicle) monetarily negate safety gains. For 20 mph versus 30 mph roads, five associational methods across 14 contexts show that the aggregate difference in the fatal or serious (KSI) collision share is negligible (0.6 percentage points). Controlling the false discovery rate, four contexts display robustly higher KSI shares on 20 mph roads: A-roads, junctions, pedestrian crossings, and single carriageways, with no context showing a robust reduction. In movement corridors where road geometry cues higher speeds, sign-only limits fail to self-enforce. These findings highlight a structural "wrong-road problem." Indiscriminate sign-only limits impose economic drag without proportionate safety returns. Transport authorities should move away from blanket statutory defaults and return to targeted physical engineering on high-harm corridors.
Danni Shi, Bryan E Shepherd, Chun Li
Tue 29 Sep
stat.ME
pdf html
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.90
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.82
- interest Semiparametric inference and debiased machine learning 0.81
Abstract
Random forests (RF) are tree-based models that capture complex, high-order predictor interactions. In biomedical studies, ordered outcomes often have outliers, skewness, or mixtures of continuous and ordinal values (e.g., due to detection limits). Standard RF splitting targets conditional means via squared errors and is sensitive to such irregularities. We propose a Wilcoxon regression tree that selects splits by maximizing a rank-based impurity reduction, equivalent to maximizing the squared Wilcoxon rank-sum statistic. Because the criterion depends only on outcome ranks, it is invariant to monotone transformations and naturally accommodates continuous, ordinal, or mixed outcomes. We develop the Wilcoxon random forest (WRF), which aggregates Wilcoxon regression trees via subsampling and estimates conditional distributions using forest-weighted empirical CDFs. We establish consistency of the WRF distribution estimator under regularity conditions. We evaluate distributional prediction using calibration diagnostics and continuous ranked probability scores, and define an out-of-bag permutation variable-importance measure. Simulations show that the WRF performs comparably to the standard quantile regression forest when errors are symmetric and homoscedastic, and yields improved quantile estimation and calibration when outcomes are skewed or heteroscedastic. We apply the WRF to predict CD4 cell count and HIV viral load six months after antiretroviral therapy initiation in a multicenter Latin American cohort, and the WRF improves threshold probability estimation and accommodates the large mass at the viral load detection limit.
Qinyou Wang
Fri 25 Sep
cs.LGstat.ML
pdf html
- favorites near “Dynamically Optimal Treatment Allocation” 0.82
- your work near “Timing as an Action: Learning When to Observe and Act” 0.78
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
Abstract
A learner's current observations need not determine its response to further training. We formulate intrinsic-extrinsic coupling through the continuation-conditioned value of a constrained learning-state intervention, with observation-relative fibers describing present agreement. An executable finite-frame classifier-head write protects current logits while repairing specified historical margins under finite-precision acceptance checks. We distinguish local admissibility, continuation-conditioned intervention value, and complete-policy performance. A matched four-cell contrast identifies readout-specific non-additivity between the same intrinsic intervention and alternative external continuations. In a CLINC-derived class-incremental setting, replay changes the write's 32-update contribution from five correct predictions to zero. Nonzero interactions also occur under output distillation, with a RoBERTa backbone, and under optimizer-native SGDW dynamics. Under SGDW, correct-count interactions are negative in all three activated roots at 128 updates, showing that coupling need not imply positive synergy. The mathematical analysis distinguishes feasible local repairs and favorable terminal outputs from training-reachable repair regions. Separate coordination tests show that content controls match or exceed the development gain, while a five-root fresh-test comparison with Fiber present in every arm shows root-dependent rather than uniformly beneficial correct-count effects. On the secondary cross-entropy readout, guided allocation yields lower mean loss than standard replay in all five pairs. Together, these results make intrinsic-extrinsic coupling operational by connecting executable state geometry to continuation-conditioned value, matched interaction identification, and closed-loop coordination, while separating identified coupling from complete-policy performance.
Weiwei Ye, Hangchen Liu, Renhe Jiang
Fri 25 Sep
stat.MLcs.AIcs.LG
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.85
- reading di, ing, versus, state
- interest Language models as statistical objects 0.82
Abstract
Large language models can interpret natural lan- guage, yet robust decisions remain challenging. Jev-like models expose structured choices, but these interfaces do not directly provide numeri- cal values at a requested precision. We propose NUMERICJEV, a training-free numerical decod- ing algorithm that enables numerical output from any LLM with a Jev-like structured-choice in- terface. Surprisingly, on our arithmetic bench- mark, it outperforms direct selection from a can- didate list containing the correct answer by 2.93 percentage points (Figure 1). Our motivation comes from the observation that numerical range selection is itself a decision problem that Jev- like LLMs can address. NUMERICJEV recur- sively refines a range through a multiway deci- sion tree while retaining the original question in context, without parameter updates or hidden- state access. On a 100-value grid, a ten-way tree requires only two decision rounds. Range- normalized MAE is 1.84% versus 5.18% for di- rect choice. A separate three-date historical- index study yields 4.58% mean relative recall er- ror and 0% readout error when the value is sup- plied. Code is available at https://github. com/Bring-AI/jev-numeric.
Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
Mon 28 Sep
cs.LGcs.AIstat.ML
pdf html
- reading dueling, style, tends, llms
- your work near “Valid Inference with Imperfect Synthetic Data” 0.79
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.80
Abstract
Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a Track-and-Stop style algorithm for best-arm identification with prescribed confidence. We prove that the algorithm almost surely achieves the asymptotically optimal cost as the error tends to zero. Finally, we extensively evaluate our approach on both synthetic and real-world instances, demonstrating consistent improvements over classical cost-unaware algorithms and their cost-aware extensions.
Shitong Wang, Zhongang Cai, Yuzhou Hong
Wed 30 Sep
stat.MLcs.CVcs.LG
pdf html
- reading open-loop, special, state, markov
- interest Macroeconometrics and time series 0.81
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.78
Abstract
Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon $K$ equals the trace of the sum of the pushed-forward innovation covariances. That error grows with $K$ after the one-step fit is exact. If the conditional mean is nonlinear, composing it is not the multi-step conditional mean. If the observation is a non-injective function of a Markov state, a memoryless one-step map does not determine future observations, while a short window can. An isotropy penalty is a function of the embedding marginal, so its partial derivative in the transition weights is zero. On a scalar autoregression with coefficient $0.9$, the one-step mean squared error is $0.998$ and the $16$-step open-loop error is $5.10$. On a hidden rotation, an eight-step window reaches $16$-step error $0.056$, while the current scalar alone reaches $0.778$. Raising the isotropy weight from $0.1$ to $10$ leaves eight-step latent error inside $[0.78,0.85]$ on three seeds.
Harshita Katoch, David Antony Selby, Gerrit Großmann, Sebastian Vollmer
Mon 28 Sep
cs.AIcs.LGstat.APstat.ME
pdf html
- interest Semiparametric inference and debiased machine learning 0.84
- reading subjective, root, fewer, neutral
- your work near “Valid Inference with Imperfect Synthetic Data” 0.78
Abstract
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.
Shuren He, Huiyan Sang, Ligang Lu
Wed 30 Sep
stat.MEstat.ML
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.83
- reading bayesian causal, nonparametric, causal, treatment
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.70
Abstract
In spatial observational studies, treatment assignment and outcomes often exhibit spatial dependence patterns, and treatment effects may vary across space and subpopulations due to both measured and unmeasured spatially structured confounders. Accounting for spatial dependence while estimating heterogeneous treatment effects (HTEs) is a central task in spatial causal inference. Causal Bayesian additive regression tree methods are popular nonparametric methods for modeling and estimating HTEs. Despite their flexibility and uncertainty quantification, the axis-aligned split rules often adopted in these models are not suitable for modeling spatial structures. We propose a spatial structure-aware Bayesian nonparametric method, called Graph-Split Bayesian Causal Forest (GSBCF), that integrates graph-split Bayesian additive regression trees (GS-BART) with the Bayesian causal forest propensity-score regression framework for spatial heterogeneous causal inference. Spatial confounding is accommodated through graph-guided split rules in modeling decision trees of the prognostic and HTE functions. We develop an efficient informed proposal sampling algorithm for posterior computation, enabling full Bayesian inference of the spatial conditional average treatment effect function. Simulations and a real data study demonstrate substantially improved estimation accuracy and uncertainty quantification over existing causal BART methods.
Yinan Cheng, Lili Zheng
Wed 30 Sep
stat.MLcs.LGstat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.84
- interest Online learning and statistical learning theory 0.83
Abstract
As black-box machine learning models become increasingly common, extracting interpretations with uncertainty quantification has become a critical challenge. One popular type of interpretation is leave-one-covariate-out (LOCO) feature importance, while prior LOCO inference methods often require data-splitting or model-refitting. A recent ensemble framework, LOCO-MP, addresses these challenges using minipatches that subsample both observations and features, but massive feature subsampling can hurt prediction in high-dimensional sparse settings. Motivated by this limitation, we consider minipatch ensembles with adaptive feature sampling guided by LOCO importance, and propose LOCO-AdaMP, which enables free LOCO inference for the resulting adaptive minipatch ensemble. We show that LOCO-AdaMP yields substantially improved predictive models while retaining asymptotically valid feature importance inference without data-splitting, despite the complex dependence between the adaptive sampling distribution and the LOCO importance statistics. Our analysis relies on a careful leave-two-out perturbation bound for the iteratively updated sampling probabilities together with the stability of LOCO scores induced by observation subsampling. Empirical results on synthetic and real datasets demonstrate advantages of LOCO-AdaMP over existing methods in predictive performance, inferential power, and stability. Overall, LOCO-AdaMP provides a flexible ensemble framework (agnostic to base models) that delivers both strong predictive performance and asymptotically valid, powerful feature importance inference for regression.
Keith Barnatchez, Willow Butler, Julia A. Geller, Elizabeth A. Stuart
Wed 30 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.83
- reading nuisance, divergence measures, functions, causal
- interest Causal inference and identification in econometrics 0.80
Abstract
Longitudinal studies with functional outcomes subject to truncation-by-death frequently report time-indexed survivor average causal effect estimates, defined among individuals who would survive regardless of treatment assignment. Because the survivor stratum can change over time, the resulting effect estimates may pertain to populations with different characteristics. We propose accompanying survivor average causal effect estimates with measures of covariate divergence between the survivor and baseline populations. We study standardized mean differences (SMDs) and the Kullback-Leibler (KL) divergence as moment-based and distributional summaries, respectively. We show that these divergence measures are identified under a subset of standard principal stratification assumptions and that their identification involves nuisance functions already estimated in survivor analyses. We derive the efficient influence functions, enabling construction of one-step bias-corrected estimators with appealing large-sample properties and allowing machine learning for nuisance estimation. We demonstrate that in randomized studies, one-step estimation of SMDs is robust to arbitrary nuisance model misspecification, whereas the KL divergence and broader density ratio estimands are more sensitive to nuisance misspecification. Additionally, we provide a sensitivity analysis framework that partially relaxes a monotonicity assumption required for identification, and construct one-step estimators for the resulting bounds. Through simulation studies, we demonstrate the finite-sample performance of the proposed estimators and illustrate how divergence measures can improve the interpretability of survivor analyses. Finally, we assess covariate divergence in survivor strata over time in amyotrophic lateral sclerosis (ALS) trials, using data from the Pooled Resource Open-Access ALS Clinical Trials database.
Udita Goswami, Shuvashree Mondal
Tue 29 Sep
stat.MEstat.AP
pdf html
- interest Semiparametric inference and debiased machine learning 0.86
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.84
- favorites near “Uniform post-selection inference for least absolute deviation regression and other Z-estimation problems” 0.89
Abstract
Linear mixed-effect panel data models are widely used in longitudinal biomedical, environmental, and social studies, but are often sensitive to data contamination and numerous covariates. To address these challenges, we propose a robust variable selection approach based on a least squares approximation (LSA) of the density power divergence (DPD) objective function combined with the Adaptive LASSO penalty. The LSA converts the nonlinear DPD objective into a computationally efficient quadratic approximation while preserving the robustness of DPD estimation. Under suitable regularity conditions, the proposed DPD Adaptive LASSO-LSA estimator is shown to possess oracle properties, including selection consistency and asymptotic normality. Simulation studies demonstrate improved robustness, better sparsity recovery, and significantly enhanced computational efficiency than penalized likelihood methods, while maintaining performance comparable to existing DPD-based approaches. An application to the SWAN bone mineral density dataset illustrates the practical relevance of the proposed methodology for robust estimation and reliable identification of important covariates.
Youssef Allouah, John Duchi
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Data Science at the Singularity” 0.84
- reading large, feedback, language model, adaptively chosen
- your work near “Valid Inference with Imperfect Synthetic Data” 0.78
Abstract
We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the number of criteria, reaching the $\Theta(\sqrt{k})$ cost of answering $k$ adaptive statistical queries with only $O(\log k)$ criteria, at fixed accuracy and confidence. In attacks on multi-task large language model benchmarks with five to ten criteria, feedback restricted to nondominated task profiles produces large reused-to-held-out score gaps and frequent false winners. These results challenge a prominent explanation for prior observed reliable benchmark reuse---that developers mainly respond to convincing improvements over the current best---in rich-feedback settings, while leaving open how often ordinary model development encounters this vulnerability.
Marthe Elisabeth Aastveit, Andreas Buschermöhle, Alex Lenkoski, Thordis Thorarinsdottir
Thu 24 Sep
stat.AP
pdf html
- reading market, markets, series, large
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.75
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.71
Abstract
Short-term vacation rentals, as promoted by platforms such as Airbnb, Homeaway, Vrbo, etc., are a growing component of the travel industry. This paper provides a unique, large dataset on global market occupancy for the short-term rental market using data from the American company Wheelhouse. The dataset consists of data for $500$ markets around the world. For each market, a daily occupancy time series from January $2017$ to December $2022$ is provided, allowing for studies of local and global patterns in the evolution of the short-term rental market. Additionally, the dataset includes curves representing the booking trajectory of each market and stay date up to one year prior to the stay date. This large dataset comprises a unique combination of time series and survival analysis data, and is suitable as a methodological benchmark for both classical statistical and machine learning models.
Troy P. Wixson, Daniel Cooley
Fri 25 Sep
stat.ME
pdf html
- favorites near “Priors for the Long Run” 0.81
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.75
- interest Macroeconometrics and time series 0.81
Abstract
Many multivariate extremes models have intractable likelihoods requiring practitioners to use alternative fitting methods. The tail pairwise dependence is a summary measure of the dependence in the tail of any multivariate regular variation model. We develop an objective function for model fitting that relies on the tail pairwise dependence as the link between our desired model (that does not have a likelihood) and a proxy model (that has a likelihood). We employ the bivariate H\"usler-Reiss distribution as the proxy model and show that there is a one-to-one relationship between the dependence parameter and the tail pairwise dependence value. Our proxy-likelihood estimator is fully developed for the transformed linear extremes time series (TLETS) models of Mhatre and Cooley (2024) and is applied to the wildfire weather data of Wixson and Cooley (2023). Simulations demonstrate that the proxy-likelihood is a competitive TPD estimator, is better at fitting TLETS models than existing methods, and is amenable to likelihood-based model selection techniques. Our estimator has smaller bias when tail dependence is weak than existing estimators reducing the need for bias adjustments. Without these adjustments, we note an increase in the tail dependence in weather-related wildfire risk between past and present climates.
Zhongze Cai, Martin B. Haugh, Xiaocheng Li
Wed 30 Sep
stat.MEstat.AP
pdf html
- reading gains, long-run, markets, appear
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.74
- interest Macroeconometrics and time series 0.79
Abstract
Scenario analysis is widely used to stress test financial portfolios, yet conventional approaches often summarize scenario gains using point estimates that overlook dependence between stressed and unstressed risk factors. We develop a machine learning framework for quantifying uncertainty in realized next-day scenario gains through prediction intervals whose widths can be calibrated online. Adaptive conformal scenario analysis (ACSA) calibrates scenario-specific quantile predictions and provides a long-run empirical coverage guarantee over realized scenarios. Our main method, kernel scenario analysis (KSA), estimates scenario-conditional quantiles directly from the specified stress and current market information. KSA can also be combined with ACSA to obtain online-calibrated prediction intervals. In experiments designed to reflect real-world markets, conventional scenario analysis can substantially understate risk, including for portfolios that appear safe under standard stress tests. Our proposed methods achieve better calibration and sharper intervals than empirical-quantile baselines. Overall, the framework moves scenario analysis beyond point estimates by quantifying predictive uncertainty and providing tools for statistically validating scenario gains.
Thomas Reiter, Christoph Kern, Fedor Miasnikov, Sofiia Nikolenko, Rob Chew, Stephanie Eckman, et al.
Wed 30 Sep
cs.CLstat.ME
pdf html
- reading basis, researchers, choice, language
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.84
- interest Language models as statistical objects 0.81
Abstract
Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive language and hate speech. Repeating the same model and task design produced high agreement (median Fleiss' $\kappa = 0.91$). Agreement fell when we changed the task design for the same tweets (median Cohen's $\kappa = 0.76$). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone. Variation across LLM task designs reached 560-572 basis points, compared with 270-331 basis points across five human instrument versions. Confidence scores did not solve this problem. They tracked repeated model outputs more closely than agreement with human labels, and grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points. We call the variation caused by task design and model choice instrument uncertainty. Researchers can measure it only by comparing reasonable task designs. Repeating one setup or relying on confidence scores cannot replace that test.
Yubin Park, Evan Brociner
Fri 25 Sep
cs.CYcs.LGstat.ML
pdf html
- reading log, dollar, little, history
- your work near “Timing as an Action: Learning When to Observe and Act” 0.80
- interest Reinforcement learning, policy learning and off-policy evaluation 0.79
Abstract
Claims data can show that provider behavior changed but cannot by itself explain why. FWA (fraud, waste, and abuse) review requires identifying material behavior, locating the codes and dollars driving it, and testing plausible explanations. A common alternative, predictive modeling, flags deviations from an expected-utilization forecast -- but a forecast has limited value unless it beats simple persistence and explains why a deviation matters. In our quarterly provider-procedure data, the latest observation captures most forecastable variation, and added model structure adds little accuracy. Residuals conflate growth, service-line shifts, code maintenance, and incomplete observation with potentially concerning behavior, making point forecasts incomplete. We instead formulate provider review as a descriptive representation problem: billed revenue y = s * p, where s measures provider scale and p describes procedure composition. The profile records scale history, effective-dated code lineage, clinical-family shares, first-use events, billing context, and Medicare-versus-client differences. An optional rank-32 nonnegative factorization of procedure co-occurrence adds a fixed semantic geometry for similarity and retrieval, surfacing evidence for review without inferring intent or adjudicating FWA. In an eight-quarter proprietary Medicare Carrier+DME audit of 1.22 million providers, simple descriptions outperform complex forecasts: regularized AR(1) attains the lowest log MAE, while persistence attains the lowest dollar WAPE. The learned semantic dictionary raises recall at 10 from 35.9% to 44.7% and high-cost-rare recall at 50 from zero to 51.8%. A lineage-aware family profile stays compact and interpretable, correlating 0.790 with learned-state movement, supporting a layered architecture where transparent descriptions form the core and learned representations add optional context.
Yanchuang Cao, Jun Liu, Tengchao Yu, Heng Yong
Mon 28 Sep
stat.MLcs.AIcs.LG
pdf html
- reading basis, gaussian process, functions, processes
- favorites near “Monte Carlo Geometry Processing” 0.80
- interest Kernels, operators and functional data 0.79
Abstract
Gaussian processes constitute a cornerstone of probabilistic machine learning, yet scaling them to large datasets typically forces a trade-off between computational efficiency and model fidelity. This work bridges this gap by presenting an adaptive multi-resolution Gaussian process framework that is both scalable and exact. Our key innovation is constructing a naturally data-sparse covariance matrix with adaptive multi-resolution basis functions. These basis functions are directly anchored to samples, eliminating the need for auxiliary points. By shrinking the support domains of multi-resolution basis, the matrix block sizes are limited, guaranteeing sparsity. The inverse of the data-sparse covariance matrix is computed exactly and efficiently via the sparse Cholesky inverse algorithm. To further improve predictive uncertainties, we construct an augmented basis function. Theoretical analysis and numerical experiments demonstrate that our model achieves exact inference with $\mathcal{O}(n \log^2 n)$ training cost and $\mathcal{O}(\log^d n)$ prediction cost, establishing a principled framework for scalable and high-fidelity Gaussian process regression.
Irene Aldridge
Tue 29 Sep
econ.EMcs.GTq-fin.GNq-fin.TR
pdf html
- reading market, makers, goods, rules
- favorites near “Price Level and Inflation Dynamics in Heterogeneous Agent Economies” 0.76
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.78
Abstract
Market microstructure studies how trading rules turn orders into prices and allocations. Those rules have been rebuilt repeatedly: for floor traders, electronic limit order books and high-frequency trading, batch auctions and dark pools, blockchains run by automated market makers and block builders, and now AI agents that discover, pay for, and compete over resources. This survey traces that evolution through one question: can a market allocate scarce goods efficiently and fairly without participants revealing everything they know and want? Each technological shift moved the binding constraint of market design from trader rationality to speed, to control over transaction ordering, to the dimensionality of what participants can report. Impossibility results persist; only their cost moves. We review evidence on high-frequency trading, trading venues, automated market makers, extractable value and algorithmic collusion, and identify allocation as the missing layer of the agent-commerce protocol stack.
Xiaotian Hou, Kwangmoon Park, Hongzhe Li
Fri 25 Sep
stat.ME
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.89
- interest Adaptive data collection and anytime-valid inference 0.80
- favorites near “Post-selection inference for causal effects after causal discovery” 0.91
Abstract
Causal discovery from observational and interventional data becomes challenging in the presence of latent confounding and selection bias, where causal structure is no longer adequately represented by directed acyclic graphs over observed variables. Existing model-free methods often rely on an exponential number of conditional independence tests and provide limited uncertainty quantification in high-dimensional settings. We develop a model-free and constraint-query optimal statistical inference framework for causal discovery under latent variables and selection using single-target interventions. We introduce the system-induced subgraph (SIS) to capture the causal relations among system variables while accounting for context variables. We establish its identifiability through maximal ancestral graphs (MAGs), and show that interventions on each observed system variable are sufficient for unique identification and necessary in the worst case. Building on these results, we develop a two-stage graph inference procedure with asymptotic family-wise error control under sufficient first-stage power. For $d_X$ observed system variables, the procedure requires at most $\frac{5}{2}d_X^2$ statistical tests, parallelizable within each stage, and achieves optimal constraint-query complexity up to a constant factor. The framework accommodates soft interventions and avoids parametric structural equation assumptions. We illustrate the methods through analysis of Perturb-seq data from interferon-$\beta$-stimulated A549 lung cancer cell lines.
Wei Biao Wu
Mon 28 Sep
stat.MLcs.LG
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.80
- reading moment, argument, gives, inequalities
- favorites near “Monte Carlo Geometry Processing” 0.79
Abstract
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison proves synchronous $L^p$ contraction when $\beta\gamma L_p<(1-\beta)(1-q_{\gamma,p})$. This direct criterion includes infinite-variance gradients for $1<p<2$, but its small-step regime requires $\beta<\mu/(\mu+L_p)$. A complementary power-Lyapunov argument establishes a positive, generally much smaller, step-size interval for every fixed $\beta<1$ and every $p>1$, using only a finite $p$th gradient moment. At $p=2$, a simpler explicit certificate gives \[ 0<\gamma<\frac{2\mu(1-\beta)^2}{L_2^2(1-\beta+2\beta^2)}. \] Its quadratic high-momentum scaling is a limitation of the chosen metric, not a sharp stability boundary. We quantify this loss, provide a general mean-only quadratic $S$-procedure, and exploit endpoint Lyapunov inequalities under stronger samplewise sector information. Verified endpoint certificates can be orders of magnitude less conservative than the explicit metric.
Frances Dean, Anna Neufeld, Joshua Barrios, Geoffrey H Tison, Ahmed Alaa
Thu 24 Sep
stat.MLcs.LG
pdf html
- reading treatment, advances, opportunity, intelligence
- your work near “Timing as an Action: Learning When to Observe and Act” 0.81
- favorites near “Stabilized Neural Prediction of Potential Outcomes in Continuous Time” 0.81
Abstract
Estimating the causal effects of medical treatments is difficult when clinically important outcomes are costly to measure or require long follow-up. Short-term or inexpensive surrogate outcomes offer a potential alternative, but surrogate biomarkers may be unavailable or difficult to identify. Advances in artificial intelligence (AI) have enabled increasingly accurate prediction of clinical outcomes from inexpensive, high-dimensional measurements, which creates an opportunity to use AI predictions themselves as surrogates. To this end, we develop a framework for estimating treatment effects from paired measurements obtained before and after treatment for each treated individual. A pretrained AI model is applied to the before and after measurements, and our estimator compares the resulting outcome predictions. We characterize the technical assumptions under which this within-person contrast identifies the average treatment effect on the treated, even when clinical outcomes are never observed for treated individuals. When these assumptions cannot be justified, we use prediction-powered inference to correct bias using a small number of observed clinical outcomes and obtain valid inference. Synthetic and real-world cardio-oncology experiments demonstrate the validity and accuracy of the approach.
Jiacheng Min, Han Li, Jean-Fran\ccois Bégin, Shuanming Li
Tue 29 Sep
stat.AP
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.76
- reading periods, climate change, payments, climate
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.74
Abstract
Climate change may materially affect long-term life insurance liabilities by altering both the level and seasonal pattern of mortality. This article develops a multi-population mortality framework that combines a Hermite spline model with a distributed lag non-linear model to capture age-, region-, and temperature-specific mortality effects. We apply the framework to mortality and temperature data from 15 Spanish NUTS-2 regions and project future mortality under three shared socioeconomic pathway (SSP) scenarios. We then assess the implications for a hypothetical whole life insurance portfolio through expected death-benefit payments and portfolio profit and loss. The results reveal an important seasonal offset: warmer conditions reduce expected payoffs during winter periods but increase them during summer periods, with these effects becoming more pronounced under more severe climate scenarios and for policies issued in later years. Over longer horizons, adverse summer mortality effects become increasingly important. The portfolio analysis further shows that climate-related mortality risk can materially increase the dispersion and downside risk of portfolio outcomes, particularly under SSP5-8.5. These findings highlight the importance of incorporating temperature-related mortality effects into long-term life insurance liability projections and risk assessment.
Se Yoon Lee
Fri 25 Sep
stat.AP
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.82
- reading procedures, curves, bart, supplementary
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.83
Abstract
Phase I cancer trials seek the maximum tolerated dose (MTD) while protecting patients from excessive toxicity. Dose assignments must therefore balance patient safety with learning the dose--toxicity relationship as data accrue. We model continuously measured toxicity outcomes using two forms of Bayesian additive regression trees (BART): isotonic BART projects posterior response curves onto nondecreasing functions, whereas monotone BART constrains the model. Joint curve and variance draws induce an MTD posterior that guides dose selection through escalation with overdose control (EWOC). We compare these methods with three parametric procedures across seven dose--toxicity curves in simulation. We assess dose-limiting toxicity (DLT) counts, above-MTD assignments, signed last-dose error, and relative absolute error. The tree procedures attained the lowest mean RAE on four nonlinear curves and jointly minimized mean DLT counts and above-MTD assignments on four curves. A Bayesian reinforcement learning perspective formulates these sequential decisions as a finite-horizon planning problem. Dose restrictions yield a lower bound on last-dose error; under exact posterior-predictive evaluation, an EWOC-based rollout policy has no greater expected weighted loss than its baseline. Dose Trial Lab, a desktop simulator for all five procedures, accompanies the supplementary material.
Jaehee Seo, Jisu Kim
Mon 28 Sep
stat.MLcs.LGmath.STstat.TH
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.81
- reading aggregate, nonparametric, manifolds, theory
- interest Language models as statistical objects 0.82
Abstract
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.
Kianté Fernandez, Xinwei Li
Tue 29 Sep
cs.LGq-bio.NCstat.COstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.81
- interest Computational methods for heterogeneous agent and dynamic economic models 0.79
- your work near “Automated Solution of Heterogeneous Agent Models” 0.80
Abstract
Amortized neural likelihoods enable computationally expensive inference for models with analytically intractable or unspecified likelihoods, but their black-box nature limits interpretability. We introduce a symbolic distillation pipeline that converts trained neural likelihoods into explicit, interpretable expressions optimized for efficient parameter estimation. Our approach uses a recovery-directed objective to guide symbolic regression toward expressions that preserve parameter-recovery accuracy rather than merely approximating the likelihood function. Candidate expressions are evaluated on held-out datasets and selected using a criterion that jointly accounts for expression complexity, parameter-recovery performance, and distributional distance from the learned likelihood. We evaluate the pipeline on the diffusion decision model, a classical cognitive model, whose analytically tractable likelihood provides ground truth for controlled evaluation. The proposed recovery-directed objective improves parameter recovery over standard symbolic-regression objectives. The resulting symbolic likelihoods enable over 100 times faster parameter evaluation than both neural likelihoods and, when available, the exact likelihood, while maintaining a manageable loss in precision. We further demonstrate these computational benefits in Bayesian hierarchical inference on empirical data. Our pipeline provides a lightweight interface for integrating symbolic distillation with existing neural-likelihood estimation methods and can be adapted to a range of simulation-based inference settings.
Songqiao Han, Kalliopi Mylona, Steven Gilmour, Matteo Borrotti
Wed 30 Sep
stat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.81
- interest Adaptive data collection and anytime-valid inference 0.83
- reading sequential, tends, large, far
Abstract
Supersaturated screening experiments study many candidate factors with few runs. The experimenter must first identify the active factors, then optimize the response. One-shot regularized regression tends to select many factors, and the second-order response surface model in the selected factors is then large and needs many further runs to fit. We propose Sequential Supersaturated Screening ($S^3$), a two-stage framework for supersaturated designs. $S^3$ builds each screening round by coordinate exchange under a new positive-cone design criterion that uses no practitioner-chosen Welch calibration constant. It then removes low-importance factors one round at a time, using a graduated quantile rule. The criterion scores a design by how far its column correlations sit above the Welch lower bound, and it adapts to the current number of runs and candidate factors. We give a round-by-round bound on the probability that a noise factor is ever fixed, under conditions on noise survival and commitment at each round, together with an explicit upper bound on the total number of Stage 1 runs. Across simulated screening problems and the Borehole benchmark, $S^3$ achieves lower Type I error and higher $F_1$ scores than one-shot cross-validated Lasso (LassoCV). It also improves optimization quality in most settings and runs faster.
Luis Razo, Eliahu Cohen
Thu 24 Sep
math.STstat.TH
pdf html
- interest Empirical Bayes, shrinkage and decision theory 0.86
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.84
- reading log, inference, gives, choice
Abstract
Many inference problems are constrained from two sides: forward information and a second constraint on outcomes. Dual Boundary Condition Inference (DBCI) multiplies them and renormalises. The rule is not new: it is the product-of-experts form and the unit-exponent member of the logarithmic-pooling family. The paper contributes a classification and reducibility characterisation. The algebra does not fix the exponent at which the inputs combine: given the statistic log p_fwd + log p_bwd, which already builds in the product form, maximum entropy supplies the family (p_fwd p_bwd)^lambda and selects no member. If the inputs are read as two equally weighted opinions, three requirements select lambda = 1/2: unanimity preservation, consistency under a common Bayesian update, and minimal symmetric Kullback-Leibler divergence. If they are read as separately applied factors, two requirements select lambda = 1: a neutral second input must leave the first unchanged, and an update to one input must pass through unchanged. For a rank-one intermediate projective measurement, the Aharonov-Bergmann-Lebowitz (ABL) rule realises this form with no extra parameter. Every positive exponent orders outcomes identically, so the choice is invisible to picking the likeliest outcome, but not to scoring a class by total mass. Under the relevant identifications, the same product form gives Bayes' rule, while a final effect proportional to the identity returns the ordinary forward Born probabilities. The second result is a Two-Boundary Reducibility Criterion: DBCI factors through the forward boundary exactly when the effective backward boundary, up to positive rescaling, does. So a fixed prior or model-fixed structural constraint adds no distinctions between cases sharing the same forward boundary, though it may encode substantial information and still materially affect the result.
Inmaculada Gutiérrez, Asier Urio-Larrea, J.Tinguaro Rodríguez, Daniel Gómez, Javier Montero, Humberto Bustince ( Faculty of Statistical Studies, et al.
Fri 25 Sep
stat.MLcs.LG
pdf html
- reading operators, partly, aggregate, theory
- interest Empirical Bayes, shrinkage and decision theory 0.85
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.82
Abstract
Aggregation theory has traditionally focused on operators defined over vectors. However, many applications-including Multi-Criteria Decision Making, Group Decision Making, Fuzzy Rule-Based Classification Systems, and overlap/grouping indices-require aggregating information naturally structured as a matrix of membership degrees (e.g., where a set of objects interacts with a family of fuzzy sets). Despite this, no formal framework has been proposed for this class of operators, partly due to the common practice of flattening matrices into vectors (which discards structural information) and partly due to a reliance on decomposable operators that aggregate rows and columns sequentially. This paper addresses this gap by formalizing the notion of a matrix aggregation operator (MAO). We analyze the decomposability and symmetry properties of MAOs, showing that certain operators cannot be expressed in decomposable form and examining several notions of symmetry. Finally, we introduce a family of MAOs termed maximum entropy global coverage indices (MEGCIs), provide a construction method for them based on combining grouping functions and MEOWA operators, and illustrate their usefulness in cluster quality assessment through an extensive computational study.
Percy S. Zhai, Ping-Shou Zhong, Wei Biao Wu
Mon 28 Sep
stat.MEmath.STstat.MLstat.TH
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.73
- interest Macroeconometrics and time series 0.81
- reading series, time series, attain, functional
Abstract
Many real-world high-dimensional time series exhibit long-memory, but Gaussian graphical model testing in this regime remains understudied. We develop a direct, data-adaptive test statistic for assessing conditional independence in the graph structure of stationary Gaussian time series. We establish a finite-sample, Berry--Esseen type Gaussian approximation bound for the statistic, which applies to both short-memory and long-memory time series. The testing procedure is fully data-adaptive using block bootstrap method, on which we provide a finite-sample validity result including in the ultra-high-dimensional scenario, and can be extended to comparing graphical structures in two-sample tests. We also develop a consistency-empowered correction to the statistic and show that such tests attain asymptotic consistency in both size and power. Our proposed method is applied to a real-world fMRI data to understand functional connectivities within brain in different periods.
Takato Yasuno, Keita Kobayashi, Ryuta Sakaguchi
Tue 29 Sep
stat.MEcs.LG
pdf html
- reading asset, markov, evidence, likelihood
- interest Bayesian computation and structural estimation 0.82
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.73
Abstract
Maximum likelihood estimation of Markov deterioration hazard models for infrastructure asset management is sensitive to the numerical conditioning of explanatory covariates, yet production pipelines often adopt a single scaling convention without systematic justification. We study four covariate scaling transformations---baseline max scaling, min-max scaling, z-score scaling, and Box-Cox transformation with a training-derived positivity shift---applied to an Exponential Hazard Markov (EHM) model estimated via L-BFGS-B on bridge inspection transition data.
Anamitra Chaudhuri, Anirban Bhattacharya, Yang Ni
Mon 28 Sep
stat.MLcs.LGstat.COstat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- interest Computational methods for heterogeneous agent and dynamic economic models 0.76
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.80
Abstract
We consider the problem of learning the underlying causal directed acyclic graph (DAG) structure corresponding to a structural equation model (SEM) with non-Gaussian errors. Motivated by an intentionally misspecified non-Gaussian SEM with all Laplace errors, we first introduce the mean absolute residual risk, defined over the space of all real matrices, and show that, asymptotically, the risk of the true weighted causal DAG matrix is strictly smaller than that of any other matrix. Nevertheless, to enhance generality and account for high-dimensional and finite-sample settings, we further incorporate row-specific sparsity penalties along with a soft DAG constraint to derive a continuous score function over the space of real matrices. Accordingly, we propose a score-based DAG learning method, named MARCEDES, formulated as an unconstrained score minimization problem, which can be efficiently solved using gradient-based optimization techniques, thereby circumventing the challenges associated with constrained optimization. Furthermore, we develop a computational algorithm to handle the non-smoothness of the score objective and to enable optimal tuning of row-specific sparsity penalties under a generalized Bayes framework. Finally, we demonstrate the efficiency and improved performance of the proposed method over existing approaches through an extensive simulation study.
Samriddha Lahiry
Fri 25 Sep
math.STmath.MPstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.84
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.81
- favorites near “Introduction to regularity structures” 0.82
Abstract
A quantum measurement transforms a quantum statistical experiment into a classical one, but different measurements generally yield experiments with different statistical information. We introduce quantum-to-classical sufficiency: relative to a prescribed measurement class, a measurement is sufficient if its induced classical experiment Blackwell-dominates those from every other measurement in the class. This separates two components of quantum inference: reducing the measurement class through symmetry or decision-theoretic arguments, and identifying a canonical measurement in the reduced class by an experiment-level sufficiency condition. We establish a factorization criterion where the parameter enters the group-averaged experiment only through weights of an orthogonal block decomposition. The block-label measurement is then sufficient relative to all invariant measurements. For tensor-product models under unitary conjugation, this identifies weak Schur sampling as the sufficient measurement for spectral inference. Combined with risk-preserving symmetrization, it reduces invariant Bayes and minimax problems over all measurements exactly to classical decision problems based on Young diagrams. We derive sharp first-order asymptotic Bayes and minimax risks for smooth spectral functionals. We establish an analogous reduction for a noncompact group action in a continuous-variable setting. We consider displaced thermal states and functionals of the thermal parameter with an unknown displacement. Within displacement-invariant measurements, total residual photon number measurement emerges as the sufficient measurement. Combined with Hunt-Stein reduction, this yields exact finite-sample reductions of Bayes-minimax and minimax problems over all measurements to classical decision problems based on the total residual photon count, which follows a negative binomial distribution.
Seungjin Choi
Thu 24 Sep
stat.MLcs.LG
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
- interest Semiparametric inference and debiased machine learning 0.84
- favorites near “Variational Bayes under Model Misspecification” 0.85
Abstract
Conformal Bayes combines Bayesian posterior predictive scores with conformal calibration, but under continuous label shift both the score and calibration weight depend on the unknown response-marginal density ratio. Existing methods typically estimate one shift parameter from pseudo-labels or predictive samples and plug it into calibration. We instead propose Joint Tilt-Sensitivity Conformal Bayes (JTS-CB), which performs sensitivity analysis over a prespecified set of plausible tilts; its split-conformal realization is JTS-SCB. Each tilt jointly determines the Bayesian conformal score and conformal importance weight. JTS-SCB forms a bounded sensitivity envelope over candidate tilts, but its calibration-only construction does not inherit the exact finite-sample weighted-conformal guarantee. We therefore study a separate candidate-weighted exact counterpart and show that its usefulness depends sharply on tail behavior. For scalar linear exponential tilts, any nonzero candidate tilt makes the exact set unbounded. More generally, tail-growing density ratios produce the same pathology, whereas quadratic tilts with a negative coefficient on \(y^2\) have vanishing tail weights and admit bounded exact inference on the original target. Ratio clipping provides a complementary bounded exact construction for a surrogate target when tails grow. Experiments show that strong plug-in predictive sampling can match the oracle when the shift is well identified, while sensitivity analysis is most useful for richer, weakly identified, or systematically biased shift models, at the cost of wider prediction sets.
William Ford
Tue 29 Sep
math.OCcs.NAmath.APmath.NAmath.STstat.TH
pdf html
- favorites near “Central Limit Theorems for Smooth Optimal Transport Maps” 0.92
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.77
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
Abstract
For probability measures on $\mathbb{R}^d$ supported in fixed compact sets, we prove that quadratic optimal transport plans are quantitatively stable in Wasserstein distance under perturbation of both marginal measures, assuming that one of the initial measures satisfies an upper Ahlfors regularity condition with exponent strictly greater than $d-1$. Under the same assumption, we also prove novel quantitative stability results for optimal transport maps. Furthermore, we prove a coercivity theorem which states that any probability measure on the product space must be quantitatively close to the set of optimal plans, if it has similar marginals and a similar transport cost to the optimal value. Finally, we prove a quantitative uniqueness theorem which acts as a quantitative counterpart to Brenier's theorem. The Wasserstein diameter of the set of optimal plans is controlled by the distance of one marginal measure to a regular measure for which uniqueness holds. In this way, "almost uniqueness" of optimal plans is quantified by the source measure being "almost regular". Examples are provided which prove that the exponents of all estimates are sharp.
Samir Orujov
Mon 28 Sep
stat.COstat.MEstat.ML
pdf html
- interest Online learning and statistical learning theory 0.85
- reading gives, theory, units, wrong
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.80
Abstract
Clusterwise least squares partitions regression data into K groups with separate linear fits. We study an ensemble whose base learner is exact: solve the problem to global optimality on each of B random subsamples of size m << n, extend each solution by nearest-surface assignment, align the labels, and combine the replicates by vote or by selection. Each replicate is then an empirical K-quantizer on m points, and the ensemble admits an exact analysis. A vote is correct at a unit once the probability of a clean subsample times the clean-data replicate accuracy exceeds one half, whatever the contaminating values; iterating the ensemble on the units it has not flagged gives a variant that estimates the trimming level rather than requiring it. With up to 20% of gross outliers in the response its worst-case accuracy was 0.89, against 0.80 for trimmed alternation at the true contamination fraction and less at every fixed level tried. Conditionally on the data the replicates are i.i.d., so the vote converges exponentially fast in B to the plurality partition of the replicate law, agreeing with the criterion minimiser outside a boundary set whose size depends on m and the micro-solver, not on B. For the two-group location model, subsamples of order 1/pi_min make a replicate right more often than wrong above a separation threshold. An O(n^3) enumeration gives the exact minimiser for two groups and one covariate, against which the theory is checked. On clean data the ensemble loses to multistart alternation at equal cost.
Yingzhen Yang
Thu 24 Sep
cs.LGcs.ITmath.ITmath.STstat.TH
pdf html
- reading inequality, gives, argument, empirical
- interest Kernels, operators and functional data 0.84
- favorites near “Operator-valued Kernels for Learning from Functional Response Data” 0.82
Abstract
We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses the modified log-Sobolev inequality for the swap walk and a two-parameter entropy closure. A peeling argument with a surrogate localization functional then gives excess-risk bounds with the same fixed-point and confidence terms as the classical inductive local Rademacher-complexity bounds, without the additional logarithmic confidence factor in earlier transductive results. For realizable learning over a binary class of VC dimension $\dVC$, with training size $m$, test size $u$, and $u\ge m\ge\dVC$, STLC yields $\cO\{\dVC\log(me/\dVC)/m\}$. This matches the standard inductive rate and, when $m\ge9$, is within a logarithmic factor of the transductive minimax lower bound of order $\dVC/m$. For transductive kernel learning, STLC gives a spectrum-adaptive excess-risk bound without the multiplicative imbalance factors appearing in the earlier local-complexity bound.
Julien Boussard, Antoine Débouchage, Théo Saulus
Wed 30 Sep
stat.MLcs.LGmath.DS
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.79
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.78
- reading equations, differential equations, structural, differential
Abstract
A wide range of methods have been proposed, including physics-informed neural networks, which are powerful but do not guarantee identifiability of the dynamics, symbolic regression, which requires a set of precomputed operations, and causal discovery, which is more principled but usually relies on strong assumptions that physical systems may violate. In this work, we develop a theory-grounded method and prove that under a set of permissive assumptions, the structural drivers and drift of stochastic delayed differential equations are identifiable. Our method outperforms others on a benchmark for driver identifiability, and on a second benchmark to evaluate physical consistency of the learned dynamics.
Junfeng Wang, Kim Zhipei Wang, Ben Van Calster, Nan van Geloven, Ewout Steyerberg, Laure Wynants
Tue 29 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.82
- reading heterogeneity, net, possible, implicit
- interest Empirical Bayes, shrinkage and decision theory 0.84
Abstract
Background Decision curve analysis (DCA) evaluates prediction model usefulness using net benefit (NB). Population-level NB is often interpreted as a proxy for population-level expected utility, but this assumes comparability of utility across subgroups. We aimed to characterize when subgroup utility heterogeneity may invalidate population-level DCA conclusions. Methods We compared prediction-driven treatment with default strategies in populations containing subgroups with different utility values. We derived an inconsistency region: combinations of subgroup-specific {\Delta}NB values for which population-level NB and utility favor different strategies. We also developed a practical robustness framework. Results Opposite signs of subgroup-specific {\Delta}NB provide a warning signal for possible inconsistency. The inconsistency region is larger when subgroup sizes are more similar and subgroup-specific \(a-c\) values are more different, where \(a-c\) is the incremental utility of a true positive relative to a false negative. When \(a-c\) differs across subgroups, population-level NB combines quantities on different implicit utility scales and may conflict with population utility. A real-world case study illustrates the problem. Conclusions Using population-level NB as a proxy for population utility implicitly assumes homogeneous \(a-c\) across subgroups. Subgroup DCA and our framework can identify and assess when utility heterogeneity may invalidate population-level conclusions.
Helena Mehler, Taulant Koka, Michael Muma
Fri 25 Sep
stat.MEeess.SP
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.86
- interest Semiparametric inference and debiased machine learning 0.83
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.81
Abstract
In genomics, imaging and clinical studies, only a few of many candidate predictors are often nonlinearly associated with a response that may be, e.g. binary, categorical, a count or a censored event time. The Terminating-Random Experiments (T-Rex) selector is a scalable variable selection method that controls the false discovery rate (FDR) by letting synthetic null variables (dummies) compete with the real predictors. While the FDR control theory embraces more general settings, to date, the T-Rex selector has been specified only for linear models. We propose a memory-efficient selection procedure with FDR control for generalized linear models and Cox regression by extending the recently developed virtual dummy construction to score-based forward selection for Bernoulli, Poisson, multinomial and Cox responses. The virtual-dummy-based selection path remains equal in distribution to explicit augmentation, so FDR control carries over under the same assumptions. Simulations confirm this equivalence and the power gained by correct model specification. Real-world applicability is illustrated on simulated genotypes and on cancer survival data.
Haixiang Sun, Jiefu Zhang, Yinghao He, Yang Xu, Vaneet Aggarwal, Bharat Bhargava, et al.
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.81
- favorites near “Dynamically Optimal Treatment Allocation” 0.83
- interest Reinforcement learning, policy learning and off-policy evaluation 0.80
Abstract
Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated responses. We introduce SPatial coHErent risk control for REplay (SPHERE), a general replay-allocation method applicable across a broad range of learning settings. SPHERE uses a representation kernel to aggregate signed prospective loss changes, attenuating unsupported spikes while retaining coherent increases. It then formulates allocation as entropy-regularized transport, redistributing uniform source mass toward supported high-risk regions while penalizing long-distance transfers. We derive replay coefficients from the transport objective's sensitivity to the original loss changes and blend them with uniform replay to maintain baseline rehearsal. Our analysis establishes conditions under which kernel aggregation improves risk estimation and bounds transport-value inflation due to residual noise and smoothing bias. Experiments demonstrate that SPHERE improves accuracy and reduces forgetting across noisy-label vision tasks, continual language-model instruction tuning, and code-generation reinforcement learning with incomplete test rewards.
Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang
Fri 25 Sep
cs.AIstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.82
- your work near “Timing as an Action: Learning When to Observe and Act” 0.77
- interest Language models as statistical objects 0.80
Abstract
Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.
Zhou Lu
Fri 25 Sep
stat.MLcs.LG
pdf html
- reading consistent, infinite, evidence, transcripts
- interest Online learning and statistical learning theory 0.86
- favorites near “The Minimax Regret of Sequential Probability Assignment, Contextual Shtarkov Sums, and Contextual Normalized Maximum Likelihood” 0.84
Abstract
When can a learner make only finitely many prediction errors along every infinite sequence labeled by a fixed, unknown hypothesis? We characterize this form of consistency for arbitrary binary hypothesis classes in ZFC, without requiring a uniform mistake bound. The characterization uses a single linear order on finite realizable traces. Each trace selects its least subtrace, and the order must satisfy two conditions: conflicting traces select different subtraces, and the order is well-founded on the traces of each fixed target. These conditions induce a learner whose selected evidence decreases on every mistake. Conversely, a consistent learner yields such an order through canonical mistake transcripts and the Kleene--Brouwer ordering. The result provides a representation of consistent prediction by finite evidence, answering a question of Lu (2024).
Peter Reinhard Hansen, Werner Ploberger, Chen Tong
Thu 24 Sep
math.STecon.EMstat.TH
pdf html
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.81
- interest Computational methods for heterogeneous agent and dynamic economic models 0.78
- reading log, eigenvalue, theory, gains
Abstract
We study identity testing for high-dimensional covariance matrices against dense alternatives of unknown direction, with $p/n \to \gamma$. Along a globally positive quadratic precision path, mixing Gaussian alternatives over a Gaussian Orthogonal Ensemble direction yields a contiguous experiment whose log likelihood reduces to the corrected Frobenius statistic; its upper-tail test attains the limiting weighted-power envelope at every fixed strength. Fixing the prior's Frobenius radius perturbs the mixture by only $O(p^{-1/2})$ in total variation, and exact whitening carries the experiment, the statistic, and its null law to any known null covariance. Separately, under a product-coordinate null, feasibility needs only $4+\eta$ moments, plus identical distributions over time when means are estimated; studentization and an exact degrees-of-freedom correction preserve the local power. A stability inequality turns near-envelope attainment into null agreement with the Frobenius rule, so uniform noninferiority on the typical dense bulk precludes gains at any contiguous alternative. For trace-matched rank-one alternatives, the corrected statistic is the first likelihood direction when $\vartheta_n \to 0$ and $n\vartheta_n \to \infty$; at fixed strength, the log likelihood ratio in the Onatski-Moreira-Hallin fixed-spike benchmark is governed by a richer linear spectral statistic below the Baik-Ben Arous-Peche threshold, while eigenvalue separation permits cost-free largest-eigenvalue enhancement above it. Simulations illustrate the theory.
Diyang Li, Fei Wang, Kyra Gan
Mon 28 Sep
cs.LGcs.DCstat.CO
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- interest Adaptive data collection and anytime-valid inference 0.82
- reading banks, institutions, fit, exchanges
Abstract
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum likelihood estimation (TMLE) refines a flexible initial fit, yielding plug-in estimators that respect the model and support efficient inference. TMLE itself, however, has remained a fully centralized procedure. To fill this gap, our paper introduces the first federated TMLE algorithm. We federate targeting itself, for an arbitrary target, loss, and fluctuation family, through two complementary frameworks. FedTMLE-G aggregates local gradients and reproduces centralized targeting step for step. FedTMLE-L lets each institution complete its own fluctuation fit before a single exchange of fitted updates, trading synchronized fidelity for local autonomy. For gradient aggregation, we develop a finite-precision protocol that transmits changes rather than values and certifies targeting accuracy within explicit bounds on exchanges and bits. A description-length analysis of the accepted updates then shows that this finite communication leaves numerical targeting error negligible against sampling uncertainty. The cost of computing an estimator is thus distinct from the complexity of selecting it. Our analysis also indicates that keeping data local is not itself a privacy guarantee of TMLE, since instability of full-record reconstruction need not prevent recovery of a specified sensitive attribute. For a personalized version of local averaging, institutions retain their own estimates and leave once local targeting is complete. A nonconvex convergence bound charges the improvement forfeited through averaging to disagreement among local fits and exposes a tradeoff between equal institutional influence and the sampling variability of small silos.
Christophe Denis, Charlotte Dion-Blanc (LPSM, CMAP), Yating Liu
Tue 29 Sep
math.STstat.TH
pdf html
- favorites near “Spectral State Space Models” 0.78
- reading nonparametric, tends, interacting, basis
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.77
Abstract
We study the supervised classification problem for interacting particle systems (IPS) from trajectory observations. The IPS belongs to one of K classes, each characterized by a distinct interaction drift. Given a learning dataset, the goal is then to predict the class label of a newly observed system based on the trajectory of a single particle. This setting raises several statistical challenges. First, the particles within each system interact and are therefore dependent. Second, although the full particle system is Markovian, the trajectory of a single particle is not. To address these difficulties, we exploit the McKean-Vlasov limit of the IPS, which describes the dynamics of a typical particle as the number of particles tends to infinity. We propose a plug-in classification procedure based on estimating the interaction drift associated with each class from discretely observed trajectory data. Hence, for each class, we provide a new nonparametric estimator of the drifts by minimizing a ridge-regularized least-squares contrast over a B-spline basis. In particular, our theoretical findings reveal that the convergence rate of the resulting classifier is of order N -1/6+$\epsilon$ for any $\epsilon$ > 0, where N denotes the number of particles in each system. Numerical experiments illustrate the performance of the proposed method and show that it outperforms an end-to-end neural network baseline that does not exploit the underlying particle-system structure.
Subhasish Basak, Anik Roy, Sourabh Bhattacharya
Thu 24 Sep
stat.MEcs.LGmath.STstat.MLstat.TH
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.82
- favorites near “Eluder Dimension and the Sample Complexity of Optimistic Exploration” 0.83
- interest Bayesian computation and structural estimation 0.82
Abstract
We propose a Bayesian nonparametric mixture of regression trees with a Dirichlet process prior over tree-parameter pairs, enabling data-driven selection of ensemble size and unifying CART, BART, random forests, and boosting. A novel splitting rule driven by the posterior predictive of a Gaussian process within each terminal node generates flexible, smooth decision boundaries; remarkably, the GP density cancels exactly in the Metropolis--Hastings ratio for GROW/PRUNE moves, ensuring computational feasibility. An exact Gibbs sampler for posterior predictive inference propagates uncertainty through random tree traversal. A parallel MPI implementation distributes independent tree updates across processors, achieving adequate speedups. We prove posterior consistency at rate $n^{-1/4}$ in Hellinger distance under only continuity of the true regression function, allowing misspecification, via the identity $h(\Theta)=0$. Simulations on Friedman benchmark show near-nominal coverage (0.94 Gaussian, 0.92 Cauchy), robust to high-dimensional noise and heavy tails, outperforming BART and bagged CART. Applications to QSAR toxicity, crime, riboflavin, wheat genomics, and air quality confirm reliable credible intervals and automatic sparsity. The DP mixture offers a principled, robust, theoretically justified alternative for challenging regression with honest uncertainty quantification.
Ryan M. Engel, Kibaek Lee, Namid Stillman
Thu 24 Sep
q-fin.RMcs.CEq-fin.CPstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.79
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.72
- your work near “Automated Solution of Heterogeneous Agent Models” 0.81
Abstract
Financial returns are heavy-tailed, and accurate tail risk estimation is central to portfolio risk management. Modern neural generators sample by pushing a simple base distribution through a learned map, and for training stability that map is built from Lipschitz components. This is the binding constraint: a Lipschitz map of a Gaussian is sub-Gaussian, so heavier-tailed targets admit no exact match at any finite Lipschitz constant. The Monge--Amp\`ere equation ties the Brenier map's local distortion to the density ratio $f/(g\circ T)$, so a deeper trough in the target density requires a higher-gain map and yields a higher-variance estimator. The argument needs only bounded distortion, so it covers normalizing flows, flow matching, GANs, and diffusion samplers alike. Semi-Discrete Optimal Transport (SDOT) relaxes the map's regularity rather than the source's tail class. Its power diagram gives every training observation a cell holding exactly $1/N$ of the source measure, and tail observations are reached by crossing a cell boundary rather than by stretching. Our primary experiment sweeps severity over a calibrated Merton jump-diffusion spanning kurtosis 94 to 1,679. SDOT holds tail ratios at $0.85$--$0.94$ with cross-seed standard deviations below $0.025$, while every learned generator either compresses the tails or inflates them with a variance that grows alongside. Further experiments carry the result to real S\&P~500 returns and to a 21-year backtest, where SDOT gives the best risk-adjusted market-neutral strategy under CVaR optimization (Sharpe $0.70$, max drawdown $-2.60\%$, against $0.40$ for the next-best generator).
Onur Batin Doğan, Fatma Sevinç Kurnaz
Tue 29 Sep
stat.AP
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.78
- reading market, price, asset, pairs
- interest Macroeconometrics and time series 0.77
Abstract
The cryptocurrency market's volatility and complex price dynamics challenge portfolio diversification and risk management. This study examines price similarity across 347 cryptocurrencies from March 2020 to April 2024 using daily OHLCV data from the Binance API. Daily dynamics are summarized via a signed, capped intraday price-range metric, and pair similarity is assessed using the Pearson cross-correlation coefficient, capturing linear co-movement at a given lag, and Dynamic Time Warping (DTW) distance, capturing shape similarity independent of lag. Data are segmented into three market phases: an uptrend, a decline, and a subsequent uptrend. For each period, the 10 pairs with highest cross-correlation and lowest DTW distance are identified. BNB-CAKE, XLM-XRP, and MANA-SAND show the strongest cross correlations across the three periods, with coefficients of 0.78 or higher and maximum correlations at lag 0, indicating synchronized movements. Three to four of the ten lowest-DTW pairs coincide with the cross-correlation top 10 per period, showing the measures are related but not interchangeable. Two same-platform fan tokens rank among the most DTW-similar pairs despite weak cross-correlation, while a euro pegged asset and gold-backed token show the lowest DTW distance, driven by low volatility rather than shared dynamics. Results demonstrate substantial, measure dependent interconnectedness with implications for diversification strategies.
Takashi Furuya, Maarten V. de Hoop, Gabriel Peyré
Tue 29 Sep
cs.LGmath.OCstat.ML
pdf html
- reading empirical, causal, sequences, relate
- interest Language models as statistical objects 0.81
- favorites near “Learning Scalable Deep Kernels with Recurrent Structure” 0.80
Abstract
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalized horizon. To relate sampling resolutions, we model tokens by $\alpha$-H\"older sequences or, more generally, a common modulus of continuity. Our notion of continuity across resolutions characterizes the causal families admitting uniform approximation on these compact input classes by a single transformer with length-independent parameters. The result extends to the infinite-length mean-field limit, where tokens form continuous curves and masked attention becomes a causal time integral. For bounded regression with target maps satisfying a $\beta$-smooth stability condition defined using regular test functions, quantitative approximation yields a generalization bound: exact empirical risk minimization over suitably sized bounded-weight transformers gives root mean-square prediction error $O((\log\log N/\log N)^{\beta/(d+2)})$ from $N$ iid labeled sequences. The bound holds at fixed confidence on the same sampling distribution, with $d$ the token dimension and no maximum-length factor. Finally, experiments on physical time series support the H\"older-regular token model at observed scales, with dataset-dependent fitted exponents, whereas text input embeddings provide a contrasting case. Native and dense sampling, shuffled controls, and refinement checks delimit this empirical regularity regime.
Chandramauli Chakraborty, Cong Ma
Tue 29 Sep
stat.MLcs.ITcs.LGmath.ITmath.STstat.TH
pdf html
- favorites near “Preference-based Online Learning with Dueling Bandits: A Survey” 0.85
- your work near “Timing as an Action: Learning When to Observe and Act” 0.79
- interest Reinforcement learning, policy learning and off-policy evaluation 0.80
Abstract
Many generative systems return multiple candidate responses and are evaluated according to the best one. Recent work shows that, when demonstrations are optimal, pass@$k$ can reduce the sample complexity of learning from demonstrations by a logarithmic factor in $k$. We ask what happens when the demonstrator is not assumed to be optimal. We find that multiple responses provide a qualitatively stronger benefit in this setting. In a finite reward-class model with no reward feedback, moving from pass@$1$ to any pass@$k$ with $k\ge2$ changes the worst-case dependence on target accuracy from $1/\varepsilon^2$ to $1/\varepsilon$, uniformly over demonstrator quality. Under standard evaluation, where an unknown reward is fixed before training, increasing $k$ provides an additional and distinct benefit: the optimal dependence on a reward class of size $N$ improves from $\log N$ to $\log N/\log k$. We further show that these two effects can be separated. Under robust evaluation, where one learned policy must compete with the demonstrator simultaneously for every reward in the class, the fast $1/\varepsilon$ dependence persists, while the $1/\log k$ improvement can disappear. We establish matching upper and lower bounds in the corresponding regimes and give a greedy multiplicative-weights learner achieving the upper bounds without any assumption on demonstrator quality.
Yiqiao Jin, Mengxin Yu
Tue 29 Sep
stat.MEmath.STstat.MLstat.TH
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.85
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.83
- interest Language models as statistical objects 0.79
Abstract
We study entity ranking and inference under the Bradley--Terry--Luce (BTL) model using pairwise comparisons collected over sparse comparison graphs. In practice, collecting high-quality human judgments can be expensive and time-consuming, motivating the use of synthetic data augmentation in model training. We analyze an iterative synthetic augmentation workflow in which synthetic comparisons generated from fitted models are successively added to the original dataset. In this process, the proportion of real data may vanish as the number of iterations grows. However, emerging literature has shown that recursive training on synthetic data can lead to model collapse, raising concerns about the statistical reliability of such augmentation procedures. To this end, we systematically analyze the resulting MLE in finite-sample, high-dimensional regimes. For the resulting iterative maximum likelihood estimator (MLE), we derive its optimal finite-sample $\ell_2$ and $\ell_{\infty}$ statistical rates and establish its asymptotic normality under natural identifiability conditions. We further characterize regimes in which model collapse is avoided despite the diminishing fraction of real data. We validate our theoretical findings through large-scale numerical experiments and an application to the Arena Human Preference 140k dataset.
Tianle Liu
Fri 25 Sep
math.STmath.PRstat.MEstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.85
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.80
- favorites near “Monte Carlo Geometry Processing” 0.79
Abstract
Heavy-tailed $p$-value combination tests are attractive under unknown dependence because their null tails can be first-order robust even when the exact dependent null distribution is unavailable. That robustness does not resolve calibration: $\Pr\{T>q(\alpha)\}=\alpha+o(\alpha)$ neither quantifies the remaining size error nor determines its sign. We develop a second-order calibration theory for positive Half-Cauchy and reciprocal, or harmonic-mean, aggregation. After exact marginal standardization, extremal dependence is represented by an index-one exponent measure on the coordinate-face lattice. Its support separates axial, full-interior, proper-face, and mixed geometries. A M\"obius decomposition shows how the noncompact weighted half-space probes every face, while quantitative face limits determine whether the correction is integrable, critically amplified, or nonintegrable. With common-factor inversion and a tail-to-calibration map, this yields machinery applicable beyond individual parametric copula families. The results include an integrable hidden-face transfer theorem, a sharp fixed-dimensional local-corner theorem with explicit higher-face control, a common-heavy-factor theorem, and size and critical-value expansions relative to independence. Gaussian, standard multivariate-$t$, positive Clayton, and max-linear pair-shock models exhibit distinct power, logarithmic, radial--angular, and proper-face mechanisms; the Gaussian weighted-half-space transfer is conditional on explicit face-boundary hypotheses. Thus dependence geometry determines the rate, coefficient, and direction of the calibration error left unresolved by first-order validity. Asymptotics use $t\to\infty$, equivalently vanishing significance levels.
Montserrat Fuentes
Thu 24 Sep
stat.MEstat.CO
pdf html
- reading theory, units, spectrum, large
- interest Kernels, operators and functional data 0.82
- favorites near “Hypoelliptic diffusion maps I: tangent bundles” 0.84
Abstract
Many scientific problems are relational: the quantity of interest is a connection between two objects, while information about similarity is available for the objects themselves. We develop a covariance framework for unordered relationships that transfers object-level geometry to the relations they form while preserving endpoint identity and invariance to ordering. Building on symmetric pairwise-kernel representations, we develop theory for the loop-free domains used in undirected networks. We establish spectral interlacing and trace-loss results after self-pairs are removed, connect the relational spectrum to regularization and risk, derive an exact inferential error for spectral truncation, and quantify how perturbations of the underlying geometry propagate to pair covariance and estimation. Simulations show when structured borrowing improves estimation and how geometric misspecification can erode that benefit. We apply the framework to autism neuroimaging using resting-state functional-connectivity data from the Autism Brain Imaging Data Exchange (ABIDE). With a 116-region parcellation and 6,670 unique connections, the application shows that a large connectome can have a much smaller effective covariance dimension. It also demonstrates that high explained covariance alone is insufficient for choosing a low-rank representation when inferential accuracy is the goal. The framework provides a principled foundation for covariance and regularization when the statistical units are unordered relationships.
Esmail Abdul Fattah, Hatem Ltaief, Håvard Rue, David E. Keyes
Fri 25 Sep
cs.PFstat.CO
pdf html
- favorites near “Monte Carlo Geometry Processing” 0.85
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
- reading matrices, gains, julia, technique
Abstract
Sparse direct Cholesky solvers fix one data structure for an entire matrix, but symmetric positive definite systems range from nearly dense to irregular, sometimes mixing both within one matrix. We let the data structure follow the sparsity structure, across matrices and across tiles within a matrix. Before numerical work starts, a lightweight selector captures the sparsity pattern of the Cholesky factor and routes the matrix, on one static shared-memory schedule, to one of three regimes: dense, sparse, or an intermediate semisparse regime. Dense tiles are stored in full; the semisparse regime uses a new active-column tile that keeps only the columns the factorization will touch, so banded or arrowhead-shaped structures common in spatial models still reach BLAS-3 efficiency. The approach suits the integrated nested Laplace approximation (INLA): one sparsity pattern factorized thousands of times with different values, so the one-time analysis cost is amortized and every factorization saving compounds. We evaluate the technique on 60 SPD matrices, comparing against MUMPS, PaStiX, CHOLMOD, symPACK, and Intel oneMKL PARDISO on Intel Xeon and AMD EPYC nodes; the selector alone achieves the lowest total factorization time in every regime. Summed over the suite, it beats the best fixed single-structure mode by 1.6 to 2.6x, and every alternative by 1.8 to 12.5x on Intel and 2.6 to 10.1x on AMD, with the largest gains on the most expensive factorizations. It trades more one-time analysis for less time per factorization, pulling ahead by the third factorization of a given pattern. As a first GPU extension, the dense route on one NVIDIA A100, with the factor resident on the device, runs 1.2 to 6.3x faster than on the faster CPU node, the margin widening with factor size. Solver, Python/R/Julia interfaces, benchmark suite, and results are open at https://github.com/esmail-abdulfattah/sTiles.
Thomas Schmelzer
Wed 30 Sep
q-fin.CPstat.ME
pdf html
- reading lasso, invested, reach, handles
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.72
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
Abstract
In a recent paper, Schmelzer and Hastie argue that Markowitz's Critical Line Algorithm and the LASSO path trace the same curve. Here we use that identity to compute efficient frontiers with a stock LASSO solver, \texttt{lars\_path} from \texttt{scikit-learn}. It handles long--short portfolios under a leverage cap, fixed leverage with varying risk appetite, and the classical long-only, fully invested frontier. Called naively, the last path stops at the maximum-Sharpe portfolio. One shift of the response, by an amount computed in advance, lets a single call reach the minimum-variance portfolio.
David Schnurr, Felix Sarnthein, Thomas Hofmann, Imanol Schlag
Wed 30 Sep
cs.LGstat.ML
pdf html
- reading pretraining, causal, state, large
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.81
- interest Language models as statistical objects 0.80
Abstract
Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large datasets. Existing linear-time alternatives, however, are mostly causal, and their potential for tabular in-context learning (ICL) remains underexplored. To address this, we (1) revisit causal training setups, (2) compare linear sequence mixers, and (3) investigate their ICL generalisation beyond the pretraining context length. First, we show that the best training setup for causal models resembles next-token prediction. Then, perhaps surprisingly, the most promising linear sequence mixer is causal: DeltaNet outperforms even non-causal linear attention. However, it degrades beyond $2$-$4\times$ the pretraining context length, and existing mitigation strategies such as bidirectionality defer the problem at best. A hidden-state oracle shows that this is not a capacity problem. Instead, our analysis points to an instability in the recurrent state, which drifts in deeper layers of causal models. Since DeltaNet's learned write rates overfit to the pretraining regime, we modulate them with a time-dependent decay schedule intervention to stabilise length generalisation. Finally, re-introducing non-causality by reading out from the final state allows us to closely match a controlled softmax attention baseline on OpenML-CC18 and TabArena.
Taiki Miyagawa
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Why Does Deep Learning Work? - A Perspective From Group Theory” 0.84
- interest Computational methods for heterogeneous agent and dynamic economic models 0.78
- reading generalization, boundedness assumptions, theory, argument
Abstract
This paper develops a theorem-level pipeline in symmetry-quotient settings: quotient linear stability implies quotient flatness, quotient flatness implies input smoothness, and input smoothness yields generalization under local covering assumptions. Flatness is often associated with generalization, and Stochastic Gradient Descent (SGD) is frequently viewed as implicitly biased toward flat solutions. However, standard flatness measures are typically defined in the raw parameter space and are therefore not invariant under function-preserving symmetries such as positive rescaling. We develop a symmetry-aware theory of quotient flatness, quotient linear stability, input smoothness, and generalization on quotient spaces of neural-network parameters. For square loss and models equipped with function-preserving group actions, we define quotient flatness as the trace of the Hessian of the empirical loss on the regular quotient manifold. We show that quotient flatness controls input smoothness through a quotient-space analogue of the flatness-to-smoothness argument. We also prove that one-step mean-square quotient linear stability of the linearized SGD dynamics implies an explicit quotient-flatness bound in terms of the batch size and learning rate, and extend this analysis to higher-order tensor moments. Finally, under local covering and boundedness assumptions, we derive population generalization bounds in terms of quotient flatness and, consequently, in terms of quotient linear stability.
Xin Xiong, Zijian Guo, Tianxi Cai
Tue 29 Sep
stat.MEstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.83
- reading institutions, future, double, year
- interest Reinforcement learning, policy learning and off-policy evaluation 0.77
Abstract
Artificial intelligence models deployed as clinical decision support tools often suffer substantial performance degradation over time and across institutions due to covariate shift, concept drift, and cross-system heterogeneity. Fully retraining complex models is frequently infeasible, particularly in EHR settings where labeled outcome data are scarce and regulatory constraints limit model modification. We propose GUARD (Guided and Uncertainty-Aware Robustness to Domain shift), a unified statistical framework for updating an existing deployed model using limited labeled target data and multiple heterogeneous source populations while guarding against future distributional shifts. GUARD formulates robust multi-source transfer learning as a semi-supervised, adversarial optimization problem anchored at the current model, and employs double machine learning with cross-fitting to obtain debiased, efficient estimates that leverage abundant unlabeled target covariates. It then constructs an uncertainty-aware, guided group DRO estimator that combines source and target information while accounting for sampling variability in the source models. Our framework jointly provides principled robustness to domain shift, efficient semi-supervised estimation, and valid statistical inference for multi-source transfer of clinical prediction models. We demonstrate the utility of GUARD via extensive simulation experiments and a real world application to predicting future disease activity in rheumatoid arthritis from EHR data spanning more than a decade. GUARD recalibrates a pretrained model with only a small number of labeled records per year and remains accurate over multi-year horizons where target-only and standard transfer estimators degrade sharply.
Yangwen Zhang, Shiwei Ni, Xiaoping Zhang, Xiaofei Guan, Lili Ju
Tue 29 Sep
stat.MLcs.LGcs.NAmath.NA
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.81
- interest Bayesian computation and structural estimation 0.83
- favorites near “A survey of sequential Monte Carlo methods for economics and finance” 0.84
Abstract
Bayesian state estimation for high-dimensional nonlinear dynamical systems entails a fundamental tension between statistical fidelity and computational tractability, as particle weights can collapse, while Gaussian ensemble updates can miss non-Gaussian posterior structure. Score-based diffusion filters offer a sampling-based alternative, but existing training-free score filters often rely on heuristic likelihood corrections, which can compromise posterior accuracy by neglecting uncertainty about the system state associated with each noisy reverse particle. To address these issues, we propose AECSF, a training-free adaptive ensemble conditional score filter. AECSF constructs an analytically tractable score estimator from the conditional Tweedie identity, which recasts noisy posterior score estimation as estimating the conditional mean of the system state given a noisy reverse particle and the observation. To estimate these conditional means efficiently, AECSF employs a shared adaptive weighted proposal ensemble, while particle-specific conditional weights yield an estimate for each noisy reverse particle without separate proposal sampling. The proposal ensemble is updated using reverse-particle information within the same reverse-diffusion run to improve conditional-mean estimation. Theoretically, we characterize when a fixed weighted proposal measure yields the exact noisy posterior score. Under stated assumptions, we establish a bound relating conditional-mean estimation errors to reverse-sampling endpoint error. Numerical experiments demonstrate that AECSF improves the accuracy of posterior sampling and nonlinear filtering in high-dimensional problems with limited forecast ensembles.
Danhyang Lee, Shinyoung Jeon, Shu Yang
Wed 30 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
- interest Causal inference and identification in econometrics 0.81
- favorites near “Deep Learning for Individual Heterogeneity” 0.78
Abstract
Transporting an outcome relationship from a source population to a target population where the outcome is unobserved requires the conditional outcome law to be stable across populations. When it is not (outcome-model drift), the target mean is not point-identified and existing transportability estimators are biased. We study settings, common in cross-cohort educational and biomedical data integration, in which bridge outcomes are recorded in both populations and can be expected to drift alongside the target outcome. The bridge splits the drift into a component it can detect, fixed by a bridge-matching equation, and a residual orthogonal to the bridge that no observed quantity restricts. Because the target is a mean, the residual acts through a single scalar sensitivity parameter $\kappa$. We prove that the identified set has a closed-form irreducible core, whose width is governed by the residual bound and by how much of the outcome the bridge leaves unexplained, and that its center is first-order invariant to the residual. These results hold for any working exponential family, with point identification arising only as the $\kappa=0$ benchmark. We develop a debiased, drift-augmented estimator that is semiparametrically efficient at anchored sensitivity, doubly robust conditional on the bridge-identified drift, and yields rate-robust Imbens-Manski inference for the set. Unlike a bridge-blind sensitivity analysis, which must bound the entire drift channel, the proposed analysis bounds only its bridge-orthogonal part. Simulations confirm that the drift estimator is unbiased under co-drift, that the core width follows its closed form, and that the set fails visibly once the residual bound is exceeded. The motivating setting is the integration of two survey cohorts whose kindergarten mathematics outcome is available in only one.
Liner Xiang, Wenbo Zhang, Hengrui Cai
Wed 30 Sep
cs.AIcs.CLstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.84
- interest Language models as statistical objects 0.81
- reading off-policy, off-policy evaluation, robust off-policy, large
Abstract
Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at https://github.com/LinerXiang/OTROPE.
Giheon Seong, Seungki Hong, Sungkyu Jung
Tue 29 Sep
math.STstat.MEstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.84
- favorites near “Variational Bayes under Model Misspecification” 0.86
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.73
Abstract
Building on the James--Stein approach to leading eigenvector estimation (Goldberg and Kercheval, Proc. Natl. Acad. Sci. USA 120, e2207046120, 2023), we develop a data-adaptive augmented James--Stein shrinkage framework for estimating leading eigenvectors and eigenspaces under a generalized spiked population model, in the high-dimensional regime where the dimension $p$ and sample size $n$ grow proportionally. For each spiked eigenvector, we construct an augmented target subspace that combines auxiliary information, either from domain knowledge or prior information, with the remaining sample spiked eigenvectors. This augmentation allows information shared across the sample spiked components to be exploited while retaining a fully data-driven shrinkage rule. We show that the resulting eigenvector estimator strictly improves upon standard PCA whenever the target subspace contains nonvanishing information about the population eigenvector, while asymptotically reverting to PCA when the target is uninformative. The individual estimators further yield a nested sequence of estimators for all leading spiked eigenspaces, with analogous dominance properties. The proposed estimator also strictly improves upon the existing HDLSS-motivated shrinkage estimator in the proportional high-dimensional regime. Simulation studies demonstrate substantial finite-sample gains and robustness to misspecification of the number of spikes.
Mohammad Zeinalpour, Amir Najafi
Tue 29 Sep
stat.MLcs.AIcs.LG
pdf html
- interest Online learning and statistical learning theory 0.87
- favorites near “A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning” 0.87
- reading generalization, fit, solely, leave
Abstract
Weak-to-strong generalization (W2SG) occurs when a student trained on a teacher's predictions outperforms that teacher. We study when this happens under fully converged, ridgeless two-stage learning, with no early stopping, no explicit regularization, and no assumption that the student is more expressive than the teacher. In two-stage linear regression, a teacher is fit from $n$ labeled examples and a student is trained solely on the teacher's predictions on $m$ fresh, unlabeled inputs. Although both stages share the same hypothesis class and the same training rule, we show that the student outperforms the teacher exactly when $m$ lies in an explicit intermediate range: too few pseudo-labels leave the student without enough signal, too many let it inherit the teacher's noise. Under power-law covariance, we derive this range in closed form as a function of the spectral decay and noise level, including regimes where the improving region splits into two disjoint intervals of $m$. We then study a random-feature model in which the student has strictly more features than the teacher, and identify two regimes, again given by explicit thresholds: one where improvement occurs only for $m$ in a bounded interval, and one where it occurs only once the student width $N_S$ exceeds an explicit threshold. Both regimes are governed by a single "active-bottleneck" principle: whichever of $m$ or $N_S$ is scarcer controls how much teacher error is filtered out, while increasing the other resource only reduces estimation noise. Together, these results show that finite data and finite width can themselves regularize a two-stage learner, with no explicit mechanism doing so.
Wonmo Koo, Jaeyeong Lee, Taeseong Yoon, Heeyoung Kim
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Spectral State Space Models” 0.81
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.78
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
Abstract
Multivariate time series anomaly detection (MTAD) is crucial for ensuring the safe and reliable operation of complex systems. Many existing methods learn normal patterns by training reconstruction or forecasting models on predominantly normal data. However, a large portion of these approaches rely on deterministic models and their associated point-wise output errors for anomaly scoring. Since real-world multivariate time series are inherently stochastic due to measurement noise and intrinsic system randomness, purely error-based scores can be unreliable, as large errors may arise from benign fluctuations rather than true anomalies. Probabilistic approaches address this limitation by quantifying uncertainty in model outputs. In particular, probabilistic state-space models (PSSMs) provide a principled framework by modeling stochastic system dynamics through latent state transitions and measurement noise via emission models. Despite this advantage, existing PSSM-based MTAD methods often struggle to capture long-range temporal dependencies and inter-variable dependencies, as they typically rely on noise-sensitive recurrent architectures and lack explicit cross-variable structure modeling. To address these limitations, we propose Graph-Transformer-Enhanced Probabilistic State-Space Model (GT-PSSM), a novel PSSM-based MTAD method that tightly integrates PSSM-based probabilistic modeling of stochastic dynamics with Graph Transformer-based learning of temporal and inter-variable dependencies. By jointly modeling stochasticity, long-range temporal dependence, and variable interactions within a unified probabilistic framework, GT-PSSM enables more robust anomaly detection.
Pengjia Cui
Mon 28 Sep
cs.SIstat.OT
pdf html
- reading research, science, differentiation, expectations
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.75
- favorites near “Cause, Effect, and the Structure of the Social World” 0.79
Abstract
Communication network research explains access to information, patterns of participation, and the organization of public meaning through different relationships and observations. This integrative review connects research topics, knowledge organization, and research practices, using bibliographic analysis of a Web of Science candidate pool to guide selective reading. Three judgments emerge from the comparisons. First, the relevance of a connection depends on the task and the criterion of value: team members anticipate consulting colleagues whose expertise they recognize, while journalists distinguish monitoring sources, using their information, and citing them. Second, commonality at one level can coexist with differentiation at another: audiences share outlets while selecting different articles, and shared issue agendas accommodate different evaluations. Third, some differences remain unresolved: contrasting media-use influence findings cannot be explained simply by whether models include selection and content co-nomination. Reading the uses of homophily, transactive memory, sourcing, agenda-setting, and framing resources clarifies which expectations and observations support these judgments. Selected citation contexts also show how conceptual and measurement resources enter the same argument. The resulting agenda calls for comparisons of relationship types across group stages, source use across reporting tasks, and encountered content with recipients' interpretations. These purposive comparisons establish specific connections among literatures without estimating their prevalence or demonstrating a common causal mechanism.
Rares Dimitrie Grozavescu, Mark Girolami
Wed 30 Sep
cs.LGmath.STstat.MLstat.TH
pdf html
- interest Kernels, operators and functional data 0.83
- reading operator, empirical, search, measured
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.82
Abstract
Projection pursuit searches for a direction along which the data look least Gaussian. When the observation space contains a large Gaussian complement, the empirical objective can be minimized by a direction that carries no signal, with empirical kurtosis as low as at the truth. Sample splitting exposes rather than repairs this failure. Appending coordinates independent of the latent regime degrades the search while leaving Bayes recoverability unchanged. Restricting the search to the column space of a known forward operator removes the failure exactly on the negative-kurtosis branch. Estimating a principal subspace from the data is the alternative. In a controlled two-component model, the leading sufficient scalings differ in the gain with which the operator transmits the discriminant: $\varsigma^{-4}$ for covariance-spike estimation and $\varsigma^{-8}$ for fourth-moment search. At fixed search dimension, the measured threshold ratio collapses onto $n/p^2$ with exponent $0.156$, close to the predicted $1/8$. This is an empirically supported scaling motivated by sufficient bounds, not a proved asymptotically tight law. When the search dimension is varied, the measured exponent is $0.325$, substantially larger than $1/8$, and the tested range does not identify its functional form. The crossing location also depends on calibration and model configuration. Under a downstream excess-error criterion, the scaling largely disappears.
Irene Aldridge
Fri 25 Sep
econ.EMcs.GTcs.LGcs.MAecon.TH
pdf html
- reading log, welfare, social, unconditional
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- interest Semiparametric inference and debiased machine learning 0.82
Abstract
We study a matching mechanism where agents and objects are described by features rather than complete rankings. A single spectral projection reduces the problem to a one-dimensional sort, computable in O(N log N) time. We prove that on descaled features and preferences, our algorithm obtains the exact Nash Social Welfare (NSW) optimum within the projected space, with an unconditional utilitarian-welfare guarantee and a conditional NSW guarantee. The proposed mechanism is stable against exogenous noise but not strategy-proof; we provide an explicit profitable misreport. On an agentic AI shopping application, the diagnostics correctly anticipate both a success and a failure case. A 100-instance robustness study confirms the findings.
Mihir Dhanakshirur, Adam Ousherovitch, Ambuj Tewari
Mon 28 Sep
cs.LGstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.83
- reading policy, large, seen, assign
- interest Language models as statistical objects 0.80
Abstract
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing different probability mass on incorrect outputs. This is formalized through a learning-theoretic counterexample in which CE minimization selects a suboptimal policy. We identify that controlling the support of the learned policy can solve this problem by preventing probability mass from spreading to unsupported outputs. Since support size is non-differentiable and computationally intractable, we propose entropy-regularized cross-entropy (ER-CE), using token-level Shannon entropy as a tractable proxy. Finally, across mathematical reasoning and code-generation benchmarks, we find that entropy-regularized training consistently improves verifier accuracy over standard cross-entropy. Our results identify a simple failure mode of imitation-based post-training in verifiable tasks and provide a practical objective that is better aligned with producing correct outputs.
Peter Reinhard Hansen, Chen Tong
Thu 24 Sep
stat.CO
pdf html
- reading derivatives, references, reviewed, quadrature
- favorites near “Simplifying debiased inference via automatic differentiation and probabilistic programming” 0.85
- interest Kernels, operators and functional data 0.78
Abstract
Fast, accurate algorithms for the Faddeeva function w(z), and hence the Voigt profile K(x,a), have existed for four decades, and analytic first derivatives are available in some implementations. What the established libraries reviewed here have not provided is the full likelihood calculus of the normalized Voigt distribution: applications still commonly resort to pseudo-Voigt approximations, finite-difference derivatives, or numerical convolution for parameter inference. Because w'(z) = -2z w(z) + 2i/sqrt(pi), every derivative of the Voigt log-likelihood is an algebraic function of K and the dispersion part L(x,a) = Im w(z), from the single complex evaluation that delivers the profile. This yields the score and Hessian in closed form, and the expected Fisher information by one-dimensional quadrature of an analytic integrand; for fixed interior widths sigma, gamma > 0, the MLE of the center and both widths is consistent and asymptotically normal at rate sqrt(n), despite the distribution having no finite mean or variance, so conventional likelihood-based standard errors apply. The conditional mean of the Gaussian component given an observation is (y - mu) - gamma L/K: a redescending function that attributes moderate deviations to the Gaussian (Doppler/resolution) component and extreme ones to the Lorentzian tail. The package voigtinference (Python, NumPy/SciPy, with a cross-validated Julia companion) supplies the toolkit: score, full parameter Hessian, expected information, Newton-based unbinned maximum likelihood with boundary diagnostics, conditional component moments, and evaluation validated against high-precision references at extreme width ratios. It applies directly to unbinned non-relativistic, constant-width Breit-Wigner x Gaussian resonance fits and supplies analytic Jacobians for line-shape refinement. Companion paper: arXiv:2605.01665.
Ruifeng Shang, Shu Liu, Yuhua Zhu
Tue 29 Sep
stat.MLcs.LGmath.DS
pdf html
- favorites near “Spectral State Space Models” 0.83
- reading equation, video generation, continuous-time, consistent
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.76
Abstract
Physical systems evolve continuously in time, yet their states are typically observed only at discrete times. Generating trajectories consistent with their probability densities from such observations therefore requires capturing the continuous-time evolution rather than only learning transition mappings between consecutive observations. We propose PhiBE-Flow, a framework that directly estimates the probability velocity field induced by the stochastic differential equation (SDE) which governs this continuous-time distributional evolution. PhiBE-Flow learns from discrete observations using a model-free approach requiring neither known SDE coefficients nor score estimation. We establish convergence guarantees for the method, accounting for both time-discretization and finite-sample errors. We evaluate PhiBE-Flow on systems of increasing complexity, from controlled stochastic numerical systems to Navier--Stokes dynamics and real-world videos. Our results show that PhiBE-Flow accurately recovers probability flows of stochastic dynamics, preserves multiscale physical statistics, and improves video generation performance over representative baselines. The code is available at https://github.com/R1fe/PhiBE-Flow.
Yang Zhou, Tianyu Guan
Thu 24 Sep
stat.AP
pdf html
- reading asset, research, reviewed, questions
- favorites near “Hicks-Arrow Prices for US Federal Debt 1791-1930” 0.76
- interest Reinforcement learning, policy learning and off-policy evaluation 0.78
Abstract
National Basketball Association (NBA) teams have increasingly detailed metrics, but better predictions do not necessarily improve decisions. This critical integrative review draws on prior reviews, citation tracing, and topic searches across seven research streams: on-court action, player value, role, lineup synergy, availability, draft and development, and contracts and roster construction. An observation-state-action-decision-evaluation chain organizes the synthesis. Six decision-readiness gates guide our assessment: point-in-time validity, uncertainty, context portability, action feasibility, opportunity-set observability, and evaluation, with requirements matched to each claim. The reviewed literature is strongest in measuring and predicting individual components of a decision. Evidence is less developed at interfaces that combine components, transfer them across settings, and compare feasible actions. We outline a proposed deployment workflow, a reporting contract, and a research agenda covering player transport, role substitution, roster fragility, legal action generation, and asset valuation. The 2023 collective bargaining agreement and forthcoming 3-2-1 Draft Lottery illustrate how institutional changes generate research questions. Models should inform evaluable comparisons of feasible choices. While its effect on organizational decision quality remains an empirical question, the framework provides a diagnostic and reporting structure for matching decision claims to evidence requirements.
Babak Barazandeh
Tue 29 Sep
cs.AIcs.LGstat.ML
pdf html
- interest Kernels, operators and functional data 0.83
- reading control, intuition, price, functional
- your work near “Computational Methods for Economic Models with Function Valued States” 0.78
Abstract
The rank $r$ in LoRA is widely treated as a capacity control: a smaller rank is assumed to yield a simpler model that generalizes better. We show that, under hard per-factor norm budgets---the idealization of the weight decay and norm control used in practice---this intuition breaks down. The reason is structural: under such budgets, the updates LoRA can reach are exactly the matrices of rank at most $r$ inside a nuclear-norm ball, and every complexity and displacement functional we analyze is maximized over this set by a rank-one update---so the rank cap never binds. The consequences follow directly. The linear-readout model class we study is identical for every $r \ge 1$, its Rademacher complexity carries no dependence on $r$, and the distance the adaptation can move the source distribution obeys a rank-independent upper bound that we show is sharp. If rank does not control capacity, where does it act? We identify two places. Statistically, replacing the per-factor budgets with a joint budget on the product restores a data-dependent, rank-sensitive complexity bound---though the gain appears only for well-spread feature distributions, and the worst case remains rank-free. Spectrally, rank sets the price of adaptation: canceling the leading singular directions of the pretrained weight requires both sufficient rank and sufficient budget. We bound the smallest rank achieving a desired source--target alignment, with upper and lower bounds that match under two-sided spectral decay. Together, these results recast rank as governing which updates are reachable and what cancellation costs---not how much capacity the model has.
Sultan Amed, Tanmay Sen, Sayantan Banerjee
Thu 24 Sep
stat.MLcs.LGq-fin.ST
pdf html
- reading income, gains, aggregate, institutions
- favorites near “Aggregating Distributional Treatment Effects: A Bayesian Hierarchical Analysis of the Microcredit Literature” 0.82
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.72
Abstract
Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.
Hazar Yueksel
Wed 30 Sep
cs.LGstat.ML
pdf html
- interest Online learning and statistical learning theory 0.84
- reading briefly, tuned, favorable, inflation
- your work near “Valid Inference with Imperfect Synthetic Data” 0.77
Abstract
Closed-form recourse moves a rejected user along the unit gradient $\hat g$ of the classifier score $f$ by the promised distance $d_p=|f(x)|/\|\nabla f(x)\|$, at which the linearized score reaches zero. We ask when this one-shot step succeeds and what additional model queries change. To leading order the step ends on the favorable side exactly when the path curvature $\kappa=\hat g^\top\nabla^2 f(x)\,\hat g$ is nonnegative. Across 80 shallow models, the fraction of rejected users whose step ends there and the fraction with $\kappa\ge0$ correlate at $r=0.985$, although on Fashion-MNIST the first falls below the second by 8.2 points on average. No rule that uses only the score value and gradient can be valid for every score with path curvature bounded by $K$ without overshooting some by order $Kd_p^2/\|\nabla f(x)\|$. When the curvature is also Lipschitz and the step is short, one evaluation of $f$ at the promised point attains the minimax rate among deterministic one-query rules that know the curvature bound and its Lipschitz constant, and split-conformal calibration makes such a rule reach the first crossing or abstain with probability at least $1-\delta$. Training with an asymmetric curvature penalty lets 99-100% of paths cross within the promised step on undershoot-prone shallow data, at about 4-22 times the overshoot of symmetric penalties (Fashion-MNIST, COMPAS). Because $\kappa$ and $d_p$ depend on how the score is scaled, part of this gain can be a longer promised step, and at matched validity a smaller audit of briefly trained models finds no uniform advantage over tuned inflation. Where a per-user line search along the ray is affordable, it is exact to grid resolution and preferable.
Yizhou Liu, Sara Kangaslahti, Jeff Gore
Tue 29 Sep
cs.LGcs.AIstat.ML
pdf html
- interest Language models as statistical objects 0.85
- reading llms, functions, likely, today
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.79
Abstract
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distributions. What happens with multiple softmax functions, as in LLMs, is unclear. Here, we show through toy models that any softmax learning peaked distributions, regardless of its position in the model, can develop logit magnitudes that grow in a power law with exponent $1/3$, becoming a training bottleneck whose loss contribution decays as a power law with the same exponent $1/3$. The overall loss therefore obeys $1/3$ scaling whenever at least one softmax learns peaked distributions. We confirm that many softmax functions in LLMs learn peaked distributions and that LLM loss scaling matches this $1/3$ prediction. Moreover, logit growth dynamics reveal that attention heads, rather than the language modeling head, are the bottleneck likely driving the $1/3$ loss scaling in LLMs. Attention trying to concentrate on specific information, which is the heart of Transformers, may therefore also be the heart of the neural scaling law of training.
John Henry V. Gray, Tianjian Zhou, Benjamin A. Shaby
Wed 30 Sep
stat.AP
pdf html
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.85
- reading american, good, cover, observable
- interest Macroeconometrics and time series 0.79
Abstract
We propose a Bayesian regression tree model fit within a multilevel structure and apply it to historic wildfire data in the western United States. Sharing of information between related groups (ecoregions) combined with highly interpretable regression trees allows for better predictions and understanding of climate and land cover variables predictive of wildfires. By doing a simulation study with a range of performance metrics, we demonstrate our method produces tree posteriors most structurally similar to assumed true trees, while simultaneously achieving good out-of-sample predictive performance. Applied to a large wildfire data set, we explore variable splits within regression trees corresponding to each ecoregion in detail, taking into account known features of each location. Shared hyperparameters between trees provide highly useful understanding of both variable and split value importance in predicting wildfires among all ecoregions, with no direct parallel in comparable models. Namely, we highlight potential evaporation, temperature, and evergreen forest land cover as variables most associated with historic wildfires, with some observable patterns in split values most commonly chosen across the groups. We propose a new algorithm based on parallel tempering, conditioning on shared hyperparameters at the true posterior temperature, improving Markov chain mixing, a known bottleneck in Bayesian CART models.
Davit Gondauri
Tue 29 Sep
econ.EM
pdf html
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.83
- reading rational, equilibrium, games, game
- interest Computational methods for heterogeneous agent and dynamic economic models 0.75
Abstract
We study a fixed-normalization class of symmetric bimatrix games generated by a common kernel and show that it admits efficient approximation despite exact symmetric-equilibrium computation remaining PPAD-hard. For every rational game in the class, a single auxiliary zero-sum saddle computation yields a rational symmetric \(1/5\)-approximate Nash equilibrium in polynomial time. We also give exact recognition and unique kernel recovery, and an exact polynomial-time post-processing algorithm that minimizes regret along the segment joining the selected saddle strategies. For arbitrary rational square games, we formulate the nearest common-kernel projection as a linear program with an explicit dual certificate, obtaining a certified \((1/5+2\eta^*)\)-approximation guarantee. A separate selector result proves a sharp regret-to-uniformity constant for the full-subset construction. The results provide a tractable and certifiable structured-game regime and clarify which common-kernel reduction architectures cannot support certain fine-grained approximation-hardness objectives.
Mostafa Shams, Robert Erhardt, Staci Hepler
Thu 24 Sep
stat.MEstat.AP
pdf html
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.77
- interest Causal inference and identification in econometrics 0.80
- reading nonparametric, large, periods, reasonable
Abstract
Data observed over space and time can exhibit dependence across both dimensions, and datasets can grow very large in both the number of locations and the number of time periods. This dependence and data size mean that many statistical models cannot be fit at a reasonable computational cost as a result of dense matrix inversions, large parameter spaces, or memory and storage challenges. When the data are ordinal, this only adds to the computational complexity of model fitting. Many ordinal models rely on a latent continuous variable designed to capture dependence in a computationally efficient manner, and then partitioned according to cutoff parameters to yield the observed ordinal response data. A very common choice is a latent Gaussian distribution, which can accommodate different dependence structures and permits Gibbs sampling in a Bayesian framework. Unfortunately, this model can be overly restrictive and fails to provide the flexibility needed to capture a range of ordinal outcomes. In this paper, we demonstrate the use of Bayesian nonparametric (BNP) methods to enhance the model flexibility of a Gaussian latent model for ordinal drought data observed over space and time, with Dirichlet process priors inducing clustering among time periods within each spatial location. In this work, we model ordinal drought data separately at each spatial location while accounting for temporal dependence, but we do not model spatial dependence across locations. We show that these BNP models often outperform Bayesian parametric approaches at a reasonable computational cost.
Jiarui Qi
Thu 24 Sep
math.STmath.PRstat.TH
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.85
- interest Semiparametric inference and debiased machine learning 0.82
- reading members, choice, uniformly consistent, consistent
Abstract
A latent severity has a bounded lower tail with density of shape alpha and scale L. It is observed only through a fixed Markov kernel K that is biased and non-additive. The relative conditional spread of K diverges at the endpoint. Our sample is i.i.d. from the marginal Q alone, with no anchoring covariate or instrument. We prove a dichotomy. The shape index alpha is identifiable: for every admissible choice of the class constants, any two observationally equivalent members of a lean class share alpha, determined by a near-endpoint expansion of Q. The rate, namely L and the fixed-scale exceedance p_tau, does not survive. There exist admissible shared class constants and two members of a smaller regularity class whose observed laws coincide exactly. Across the pair alpha agrees, whereas L and p_tau move. A degenerate Le Cam two-point bound excludes any uniformly consistent estimator of either, and pointwise consistency fails at one member. Only the rate needs an anchor. We conjecture that a known kernel family with known edge map identifies the rate fiber by fiber if and only if the family satisfies a fixed-scale injectivity clause, and we prove the sufficiency direction. In surrogate safety, uncalibrated conflict data give the shape of near-crash risk, not its absolute rate.
Shucheng Liu, Chan Park, Guanhua Chen
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.83
- favorites near “Deep Learning for Individual Heterogeneity” 0.81
- reading treatment, doubly robust, doubly, estimating causal
Abstract
Estimating causal effects with continuous treatments in observational studies is challenging due to confounding, model misspecification, and high-dimensional covariates. We propose the Weighted Spline-Expanded Network (WSENet), an end-to-end neural framework that addresses these challenges by combining covariate balancing, structured treatment embedding, and bias-corrected outcome estimation. WSENet first applies Distance Covariate Optimal Weights to induce distributional independence between covariates and treatment without relying on parametric models. It then learns the conditional outcome via a structured network that fuses outcome-relevant representations of covariates with a spline-expanded treatment input, enabling smooth and flexible modeling of the dose-response relationship. To mitigate residual bias, we introduce Weighted Targeted Regularization, a correction technique based on efficient influence functions that yields a doubly robust estimator. Extensive evaluations on semi-synthetic and real-world datasets, including high-dimensional genomic and environmental health data, demonstrate that WSENet consistently outperforms existing baselines in both accuracy and stability.
Alexandre Declèves, Etienne Boursier, Nicolas Flammarion
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.85
- reading pretraining, target parameters, modern, implicit
- interest Language models as statistical objects 0.79
Abstract
Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.
Pengyu Zhang, Mark Girolami, Arnaud Vadeboncoeur
Tue 29 Sep
cs.LGstat.ML
pdf html
- reading equation, science, tradeoffs, unconditional
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.78
- interest Language models as statistical objects 0.80
Abstract
Inferring quantities of interest (QoI) from data is a central task in Science and Engineering. In such contexts, we often have access to both quantitative data and qualitative data. Quantitative data may be represented by noisy sensor measurements, simulation data, re-analysis data; qualitative data may be in the form of text descriptions of experimental setups, expected experiment outcomes, and human-perceived system behaviours. The task we address in this paper is the following. Given a training set of paired qualitative text and quantitative QoI data, we learn to exploit the inherent correlation between the two modalities to learn a highly informative data-driven natural-language-conditional Bayesian prior, such that when presented with a new physical system, we can coherently combine (i) the training dataset, (ii) qualitative text describing the new system, and (iii) a small number of noisy sensor readings from that new system, to perform inference and uncertainty quantification (UQ) over the QoI. To achieve this task, we develop two parallel approaches, one uses conditional diffusion and the other conditional autoencoders, and compare both against classical Bayesian methodology, unconditional generative models and deterministic supervised methods. Each approach has specific strengths and tradeoffs; conditional autoencoder offers theoretical tractability, allows for fast posterior sampling, and provides better-calibrated UQ, whereas conditional diffusion is explored for greater expressiveness and capturing complex posteriors with irregular QoI fields. The approach is tested on the steady-state heat equation, damped Helmholtz equation, and UK weather reanalysis data.
Yingli Wang, Kelvin Shuangjian Zhang, Lingjiong Zhu
Fri 25 Sep
math.NAcs.NAmath.OCmath.PRstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.83
- favorites near “Finding Regularized Competitive Equilibria of Heterogeneous Agent Macroeconomic Models via Reinforcement Learning” 0.78
- your work near “Automated Solution of Heterogeneous Agent Models” 0.78
Abstract
We study global convergence guarantees of third-order Langevin dynamics for non-convex optimization via simulated annealing with fixed friction and decreasing noise. An explicit three-block distorted entropy transfers dissipation from the noisy auxiliary variable to the full state. Under dissipativity, regularity, and low-temperature functional-inequality assumptions, logarithmic cooling drives the objective values to the global minimum in probability at the barrier-controlled kinetic rate. For the exact-force-integral and midpoint three-stage discretizations, polynomially decreasing steps preserve this rate on the physical time scale. The cubic local endpoint estimate gives a less restrictive sufficient step-size condition than the available frozen-force kinetic result. A comparison with the one-gradient UBU integrator shows how its centered stochastic local error leads, under the same strong-coupling analysis, to a smaller sufficient iteration exponent. Numerical experiments are conducted to illustrate our theory. For a double well objective, third-order Langevin terminal-success point estimates are higher than UBU at both a common horizon and an equal gradient budget. For a high-dimensional nonconvex neural-network objective using synthetic data, independently tuned UBU and third-order Langevin schemes both outperform overdamped Langevin dynamics; the third-order Langevin point estimate is higher. For the same neural-network objective on real data, we show the same point-estimate ordering for best-basin probability and post-quench test accuracy. Numerical code and associated experiment results are publicly available at https://github.com/gagawjbytw/simulated-annealing-third-order-langevin.
Yanfeng Yang, Junda Zhao, Yijie Gao, Jiaqi Yang, Xinyu Shi, Ziqi Chen, et al.
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.82
- interest Language models as statistical objects 0.79
- favorites near “Priors for the Long Run” 0.78
Abstract
Conditional independence (CI) is a fundamental concept in statistics and machine learning. Recent advances in conditional generative modeling provide flexible tools for generative-model-based CI tests, which rely on an estimated conditional distribution to generate randomized samples. However, errors in estimating this distribution accumulate in existing Type I error bounds, and consistency of the generative estimator alone does not guarantee asymptotic Type I error control. To address this limitation, we formulate conditional generative modeling as a domain adaptation problem and leverage auxiliary data from multiple source domains to improve estimation in the target CI testing domain. We propose Domain-Adapted Diffusion (DA-Diff), a multi-source domain adaptation framework for conditional diffusion models based on weighted empirical risk minimization over both target and source domains. We establish the convergence rate of DA-Diff and show how transferable source data can improve target-domain estimation through an increased effective sample size while controlling transfer bias. Building on DA-Diff, we further propose Domain-Adapted Conditional Independence Testing (DA-CIT) and show that its Type I error satisfies $P(p \leq \alpha) \leq \alpha + o(1)$. Experiments demonstrate that DA-Diff improved conditional generation quality compared with transfer-learning diffusion baselines, while DA-CIT provides strong Type I error control and competitive power.
Mana Sakai, Masaaki Imaizumi
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.82
- interest Kernels, operators and functional data 0.82
- reading govern, large, horizons, eigenvalues
Abstract
Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as $n,T\to\infty$ whenever $T=o(\sqrt{n})$. (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.
Taulant Koka, Michael Muma
Mon 28 Sep
eess.SPstat.ME
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.81
- interest Adaptive data collection and anytime-valid inference 0.81
- favorites near “Measuring Price Selection in Microdata: It's Not There” 0.79
Abstract
Reliable high-dimensional variable selection requires scalable error-controlling methods. The Terminating-Random Experiments (T-Rex) selector estimates the false discovery rate (FDR) by aggregating early-terminated forward-selection paths in which predictors compete with synthetic dummies. We address two remaining challenges: i) predictor dependence can bias dummy-predictor competition; ii) computation is wasted on recomputing terms shared across experiments. Building on memory-efficient virtual dummies, which sequentially sample projections from their exact conditional law, we estimate the conditional Gaussian law of inactive predictors and map draws onto the remaining sphere radius. A shared lazy Gram cache computes the response product and each requested Gram column once across experiments. Simulations across three covariance structures show that uniform spherical dummies can exceed the target FDR, whereas the proposed method empirically controls FDR. Caching yields more than a sixfold speedup, and the method remains feasible with 100000 predictors, where competing FDR-controlling methods become computationally impractical.
Wondesen Teshome Bekele, Marco D'Oria
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.81
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.80
- interest Bayesian computation and structural estimation 0.81
Abstract
Many probabilistic predictors must satisfy exact structure in every stochastic realization, yet common hard-constraint approaches form predictions in ambient coordinates and then correct or project them. We introduce SLP-ProbHard, a cross-family, representation-centered framework for probabilistic hard-constrained learning when explicit structural parameterizations are available. Its core object, a Structural Feasible Latent Parameterization (SFLP), combines a structural latent law $Z \sim P^Z_\theta(\cdot\mid x)$ with a feasible map $Y=h_\phi(x,Z)$ that satisfies the constraint for every latent realization. Together these components define the predictive law itself, including its support and boundary probabilities, rather than serving as a final feasibility wrapper. We study how feasible coordinates and maps affect stochastic dimension, dependence, calibration, expressiveness, and computation. Experiments use Gaussian latent laws and fixed geometry-derived maps across affine equalities, ordering and simplex constraints, nonlinear manifolds, and three structural representations of seven-basin hydrological flow-duration-curve (FDC) data. In an official-source affine comparison with ProbHardE2E/DPPL, both methods achieve zero practical constraint violations. SLP-ProbHard uses 8 instead of 11 stochastic coordinates and improves MSE/MAE, while DPPL yields better marginal CRPS and closer-to-nominal coverage; a paired test detects no Energy Score difference across ten seeds. Real-world affine and nonlinear FDC representations reduce 13 to 7 and 14 to 8 ambient versus computational coordinates, respectively. Exact feasibility alone thus does not determine a predictive law, motivating direct structural generation when meaningful feasible coordinates are available.
Percy S. Zhai, Maggie Cheng, Wei Biao Wu
Tue 29 Sep
math.STstat.MEstat.MLstat.TH
pdf html
- interest Macroeconometrics and time series 0.81
- reading series, consistent, theory, little
- favorites near “Advances in Nowcasting Economic Activity: Secular Trends, Large Shocks and New Data” 0.77
Abstract
Conformal prediction provides marginal coverage guarantees, yet practitioners may wonder if the observed coverage is truly abnormal or consistent with sampling variation. Inference for realized coverage has received comparatively little attention, especially for time series. We study split conformal prediction with adjacent calibration and test sets of temporally dependent data. Using the functional dependence measure, we derive non-asymptotic bounds on marginal coverage error without mixing assumptions, which can be difficult to verify and may fail even for simple short-memory models. We establish a Bahadur representation to derive, to our knowledge, the first central limit theorem for realized coverage of split conformal prediction under temporal dependence. A consistent block-based estimator of the standard error yields an asymptotically justified test. We further study long-memory time series, which remain understudied in conformal prediction. For Gaussian linear processes, we show how very strong temporal dependence can lead to a non-Gaussian limiting law of realized coverage and establish block-sampling inference with an estimated normalization. The resulting theory explains how temporal dependence changes coverage uncertainty.
Jing Li, Haibin Zhu
Fri 25 Sep
stat.MLcs.LG
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.79
- reading sequential, confidence sequences, forecasting, modern
- interest Adaptive data collection and anytime-valid inference 0.81
Abstract
Modern conformal forecasting systems often maintain several adaptive pipelines that differ in base forecasters, conformity scores, calibration windows, and update rules. Comparing them is difficult because coverage is a hard constraint, whereas efficiency should be optimized only among feasible pipelines. We formulate this problem as sequential inference for a stochastic constrained argmin. At each time, the target is the set of minimum-cost pipelines satisfying multiple prefix-average conditional miscoverage constraints. We introduce Coverage-Constrained Sequential Model Confidence Sets (CC-SMCS), which separate certifiably feasible, possibly feasible, and possibly constrained-optimal pipelines. Using simultaneous martingale confidence sequences, CC-SMCS projects a rectangular confidence region onto the constrained argmin and admits an exact closed-form rule. With probability at least $1-\delta$, it contains every constrained-optimal pipeline simultaneously over all times. This finite-sample guarantee requires no stationarity or mixing assumptions and remains valid under data-dependent stopping. We also establish an impossibility result for safe certification at the coverage boundary and extend the construction to delayed multi-horizon feedback and outcome-dependent efficiency objectives.
Kangjie Zhou, Ming Yuan
Tue 29 Sep
math.STstat.MEstat.TH
pdf html
- interest Kernels, operators and functional data 0.82
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.79
- reading empirical, theory, functional, generally
Abstract
Spectral risk measures, including conditional value-at-risk (CVaR), generate a family of convex criteria for portfolio estimation. We study how the criterion should be chosen when different criteria share the same population minimizer, and whether the efficient criterion can itself be learned from data. We develop a general asymptotic theory for empirical spectral-risk minimization under estimated linear constraints. Under normal scale-mixture elliptical returns, all spectral risk measures identify the same population efficient portfolio, but their empirical minimizers have different sampling distributions. Their asymptotic covariance decomposes into a common component and a positive-semidefinite component scaled by a functional of the spectral measure, reducing efficiency to an optimization over probability measures. We characterize the efficiency-optimal spectral measure and show that single-level CVaR is generally inefficient. We then construct a fully data-adaptive estimator that learns the radial distribution and optimal spectral measure from the same observations used for portfolio estimation, yet has the same first-order distribution as the infeasible oracle. Simulations and an empirical application illustrate the method.
Izak Schmidlkofer, Zhirui Hu, Lishan Shi, Weixiang Yu, Hyungsuk Tak
Thu 24 Sep
stat.APastro-ph.IM
pdf html
- favorites near “Priors for the Long Run” 0.82
- interest Macroeconometrics and time series 0.82
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.75
Abstract
Large-scale astronomical surveys provide unprecedented volumes of multivariate time-series observations obtained through multiple optical filters. We develop a structured multivariate continuous-time autoregressive moving average (MCARMA) framework for multi-band time series with irregular sampling, heteroscedastic measurement errors, and partially observed bands. The framework allows band-specific stochastic dynamics while modeling cross-band dependence through correlated Brownian driving processes, with state-space and spectral representations enabling likelihood-based inference and interpretation of the fitted stochastic dynamics. We develop a two-stage estimation procedure in which a numerically stabilized preliminary fit initializes subsequent maximum likelihood estimation. Simulations show that higher-order stochastic structure can be recovered when its characteristic features are adequately resolved, but can become weakly identifiable because of limited temporal resolution or near pole--zero cancellation. Joint multivariate estimation improves parameter recovery in 23 of 27 settings and spectral recovery in 26 of 27 settings relative to separate single-band fits. Three Sloan Digital Sky Survey Stripe 82 quasars, respectively favoring MCARMA(1,0), MCARMA(2,0), and MCARMA(2,1), illustrate how joint multiband modeling uses cross-band dependence to inform marginal dynamics and can yield different model-order and spectral inference. The methodology is implemented in the Python package mcarma.
Shayma Alkobaisi, Anas Ali
Tue 29 Sep
stat.APcs.AIcs.CYcs.LG
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.82
- reading counterfactual, nominal, causal, nonlinear
- interest Causal inference and identification in econometrics 0.79
Abstract
Low-cost indoor air-quality sensors could support personalized asthma prevention, but their nonlinear measurement error, delayed exposure effects, time-varying confounding, and heterogeneous lower-tail responses limit risk estimation. We present CAUSALQUANT-ASTHMA, a measurement-error-aware causal quantile distributed-lag framework for short-horizon peak expiratory flow analysis. Sparse reference measurements train a nonlinear calibration model; stabilized sequential generalized-propensity weights address measured exposure assignment; and a susceptibility-modulated, smooth, noncrossing quantile model estimates lag-specific and sustained-exposure contrasts. Because no authorized cohort simultaneously provided dense indoor sensing, reference co-location, and outcome-compatible longitudinal data, evaluation used five semi-synthetic panels with known counterfactual truth, 150 patients and 12,600 patient-days per realization. Across eight methods, CAUSALQUANT achieved a dose-response integrated absolute error of 0.304 plus or minus 0.094, improving 24.2 percent over the strongest measurement-error and propensity-weighted baseline. It also obtained the lowest overall pinball loss, 1.065, while maintaining zero quantile crossings and 78.1 percent coverage for the nominal 80 percent interval. Sensor calibration reduced held-out exposure RMSE by 33.7 percent. Stress tests quantified degradation under sensor noise, missing personal measurements, and hidden confounding. These findings establish methodological feasibility and reproducibility, not clinical effectiveness; prospective, governance-approved external validation is required before patient-level interpretation or deployment.
Sang Bin Moon, Nicole Cho, Daniel Borrajo, Sumitra Ganesh, Abolfazl Hashemi
Mon 28 Sep
cs.CLcs.AIstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.83
- reading versus, generating, language, nearly
- interest Language models as statistical objects 0.77
Abstract
Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, uniform-coverage sampling, and evolutionary persona generation. Evaluations on the Alternative Uses Task (AUT), Infinity-Chat, and Divergent Association Task (DAT) show the benefits of the proposed methods across tasks and creativity objectives. On AUT, evolutionary persona generation increases response diversity by 78.8%, originality by 26.1%, flexibility by 49.5%, and holistic creativity by 13.9% over task-only prompting, while maintaining 98.5% validity; on Infinity-Chat, it nearly doubles persona-induced response separation relative to random personas. Moreover, evolutionary personas compose with creativity-optimized prompting, further increasing its response diversity by 18.6% and creativity by 6.3%. These results establish persona-set geometry as a task-agnostic mechanism for eliciting divergent LLM outputs, and support persona diversification as a reusable complement to prompt optimization.
Qijun Tong, Masahiro Ikeda, Ryota Kawasumi
Tue 29 Sep
stat.MLcs.LG
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.79
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.77
- favorites near “Recent Developments in Machine Learning Methods for Stochastic Control and Games” 0.76
Abstract
Stochastic gradient descent (SGD) is widely used to train machine learning models, but subsampling the training data introduces noise into its updates. The strength and applicability of high-probability guarantees therefore depend critically on how the tails of gradient noise are modeled. Reports of heavy-tailed gradient noise in deep learning motivate relaxing the bounded-noise and sub-Gaussian assumptions commonly used in high-probability analyses of SGD. We use Young functions from Orlicz space theory to describe noise tails in a common framework. We model SGD gradient noise by adopting a Young function that preserves the finiteness of all polynomial moments while allowing tails heavier than sub-Weibull, including lognormal distributions. The resulting class is called $\beta$-heavy-tailed, with $\beta$ controlling the tail heaviness. We establish concentration inequalities for $\beta$-heavy-tailed noise and combine them with a uniform bound on the difference between empirical and population gradients along the SGD trajectory to obtain high-probability bounds on optimization and population-risk stationarity for smooth nonconvex losses under trajectory assumptions. The bounds are not restricted to a particular learning-rate decay rule and make explicit the effects of noise tails and learning-rate schedules. Under the Polyak-{\L}ojasiewicz condition, we bound the risk at the last iterate. We also analyze SGD with gradient clipping under the $\beta$-heavy-tailed noise model.
Yeseul Jeon, Aaron Scheffler, Rajarshi Guhaniyogi
Wed 30 Sep
stat.ME
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.79
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.71
- reading mcmc, functions, gp, variational
Abstract
Varying coefficient (VC) regression models have become indispensable tools in spatial data analysis, providing unmatched flexibility in capturing complex, nonlinear relationships and spatially-varying effects of predictors on responses. Although hierarchical Bayesian approaches offer a rigorous probabilistic framework for uncertainty quantification in VC modeling, their practical application to large-scale spatial datasets remains severely hindered by the scalability limitations of Markov chain Monte Carlo (MCMC) algorithms. In response, the past decade has seen substantial advances in developing more efficient hierarchical Bayesian VC models, primarily through substituting traditional Gaussian processes (GP) with computationally efficient stochastic surrogates for estimating unknown coefficient functions. This article introduces a fundamentally different approach: the Geostatistical Variational Auto-Encoder (GeoVAE), a hierarchical deep generative framework built specifically for joint estimation of multiple spatially varying coefficient functions. GeoVAE departs from both classical GP-based Bayesian models and standard variational auto-encoders (VAEs) in two key ways. First, it constructs a coefficient-specific auto-encoder for each coefficient function, allowing each to capture its own spatial resolution and smoothness. Second, a hierarchical synthesis layer integrates information across these auto-encoders to a shared auto-encoder, explicitly modeling cross-coefficient dependencies arising from the shared spatial domain. This design enables GeoVAE to recover complex spatial patterns at computational cost that scales favorably with sample size, without MCMC or explicit GP covariance representations. We characterize the advantages and limitations of GeoVAE relative to hierarchical Bayesian spatial models, demonstrating its strong potential for large-scale spatial analysis.
Kushaal Kumar Pothula, Arunav Kumar
Fri 25 Sep
physics.plasm-phstat.ML
pdf html
- reading operators, operator, state, grids
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.81
- interest Computational methods for heterogeneous agent and dynamic economic models 0.75
Abstract
Direct numerical simulation of magnetic reconnection is limited by the scale separation of resistive magnetohydrodynamics. At large Lundquist numbers $S$ the current layer thins as $S^{-1/2}$, forcing fine grids and short time steps that make parameter scans prohibitively expensive. Deep learning neural operators offer an alternative by learning the map between function spaces rather than individual solutions, so that a single trained model returns the state for any parameter and time at inference cost. We have developed a Fourier Neural Operator (FNO) based Physics-Informed Neural Operator (PINO) surrogate for two-dimensional compressible, viscous, resistive reconnection in a wall bounded domain. We condition on the initial state, the Lundquist number, and a continuous query time. Predicting the magnetic flux function makes $\grad\!\cdot\!\bB=0$ exact, direct time queries remove autoregressive error accumulation, and parity-aware spectral differentiation matches the wall treatment of the reference solver. Trained across $10^{3}\le S\le2\times10^{5}$ with two values withheld, the surrogate recovers density, pressure, and guide field to better than $0.2\%$, velocity to $3.2$ to $3.9\%$, the flux function to $1.0\%$, and the spectral current density to $7.5\%$. It resolves the thin, narrow current sheet, Alfv\'enic jets, and spectral resolution beyond mode cutoff. Our surrogate reproduces the layer-averaged reconnection rate to within $3\%$ through $S=1.1\times10^{4}$, recovering a reconnection time exponent of $0.497$ consistent with Sweet-Parker scaling. When queried zero-shot on the $2049^{2}$ solver mesh, a full trajectory run is on the order of seconds on one GPU against several hours for the DNS, over two orders of magnitude faster, establishing learned operators as a practical route to surveying reconnection across parameter space.
Ann Huang, Mitchell Ostrow, Zhouyang Lu, William T. Redman, Leo Kozachkov, Kanaka Rajan
Wed 30 Sep
cs.LGq-bio.NCstat.ML
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.84
- interest Computational methods for heterogeneous agent and dynamic economic models 0.78
- your work near “Timing as an Action: Learning When to Observe and Act” 0.74
Abstract
On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by low-loss regions. Yet how their internal computation varies within these regions is unknown. A parallel line of work has identified the degeneracy of neural representations: many networks reach similar training loss with distinct internal structures. However, it is unclear how these solutions are related in weight space. We unify these subfields and show for the first time that many different internal mechanisms exist within a local mode-connected region in weight space. To do so, we introduce Hessian Null Space Continuation (HNC), a scalable method that uses local curvature to traverse regions of weight space that preserve network function, and can be steered toward solutions with specified properties. In RNNs trained on a memory task, HNC reaches drastically different representations and dynamics with maintained behavior. In ImageNet-trained Vision Transformers, HNC finds representations that differ more from the original network than any independently trained model with a different architecture or objective. In reinforcement-learning agents, HNC uncovers a distinct navigation strategy at comparable return and exposes reward hacking in an AI Safety Gridworld. Finally, HNC measures the local geometry of the solution set, showing how model size and task complexity shape its dimension and functional sensitivity. Our results show that a surprisingly large amount of representational diversity exists near a single trained solution, unseen by standard gradient-based optimization. HNC identifies and quantifies this diversity, opening new possibilities for mechanistic understanding of solution spaces and for model merging, editing, and fine-tuning.
Sunwoo Lim, Mladen Kolar
Mon 28 Sep
stat.ME
pdf html
- interest Semiparametric inference and debiased machine learning 0.85
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.79
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.79
Abstract
Multitask linear models can improve estimation and prediction by exploiting structure shared across responses, bridging taskwise fitting and complete pooling. In many applications, however, the objective is estimation in a data-limited target domain, while data-rich but heterogeneous source domains are available. Borrowing from these sources can improve efficiency but introduce bias. We develop a joint-sparse transfer-learning framework for high-dimensional multi-output regression that combines shared predictor structure across responses with source-target similarity. The framework yields two complementary estimators: a fused estimator that aggregates jointly fitted domain-specific coefficients and a target-based debiased estimator that adjusts for source-induced shifts. Our error bounds show how transfer increases the available information and sharing predictors across responses reduces selection costs. They also reveal a tradeoff: the fused estimator benefits from larger sources but may retain bias if source shifts point in similar directions, whereas debiasing trades this bias for additional estimation error governed by the smaller target sample. Comparison with a minimax lower bound identifies regimes where the bounds match up to logarithmic factors, where matching remains unresolved, and where projection onto a target-based convex set closes the gap. Simulations and an analysis of single-cell RNA and surface-protein profiles across cell types support the theory.
Marcus Rockel
Thu 24 Sep
math.STmath.PRstat.TH
pdf html
- interest Semiparametric inference and debiased machine learning 0.86
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.72
- reading inequality, price, given, value
Abstract
Blest's rank correlation $\nu$ is a variant of Spearman's rho $\rho$ that weights the leading ranks of one variable more heavily, at the price that $\nu$ is not symmetric in its arguments. We quantify both features by determining the exact region of $(\rho,\nu)$ over all bivariate copulas, as well as that of $(\eta,\nu)$, where $\eta$ is the symmetrized Blest coefficient of Genest and Plante. The latter region is a linear image of the set of all pairs $(\nu(C),\nu(C^\top))$ formed by a copula $C$ and its transpose. Consequently, Blest's coefficient differs from Spearman's rho by at most $1/4$, and interchanging the two variables changes it by at most $27/64$, improving on the bound $1/2$ implied by the first inequality. For every given value of $\rho$ or $\eta$, each corresponding extreme value of $\nu$ is attained by exactly one copula, given in closed form and supported on finitely many line segments. Near countermonotonicity, the upper extremizers of the $(\eta,\nu)$-region are supported on the graph of a function of the second coordinate, yet their conditional laws given the first coordinate carry two atoms. The proofs rest on a rearrangement inequality with equality case and on explicit Kantorovich potentials.
Arthur Paing, Arthur Jacot
Tue 29 Sep
math.OCcs.LGstat.ML
pdf html
- interest Online learning and statistical learning theory 0.85
- your work near “Computational Methods for Economic Models with Function Valued States” 0.82
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.81
Abstract
This paper studies convex optimization when the gradient cannot be evaluated exactly, but only approximated by a hierarchy of algorithms whose compute grows like $\delta^{-\gamma}$ in the accuracy $\delta$. When $\gamma>2$, falling into the Harder-Than-Monte-Carlo (HTMC) regime, the price of accuracy outruns the variance reduction that Monte Carlo would buy and we show that minimizing a loss function costs no more, up to a factor depending only on $\gamma$, than a single evaluation of its gradient at the accuracy the problem demands. A randomized multilevel oracle replaces the deterministic approximation of accuracy $\delta$ by an unbiased estimator of it, whose variance $\sigma^2$ becomes a second, independently priced dial: the cost of one call drops from $\delta^{-\gamma}$ to $\delta^{2-\gamma}\sigma^{-2}$. Plain inexact gradient descent driven by that oracle reaches loss $\varepsilon$ at expected compute $\Theta(\varepsilon^{-\gamma})$ in the convex case, against $\Theta(\varepsilon^{-(\gamma+1)})$ for the same method run at a fixed accuracy: randomization buys a full power of $\varepsilon$. Under $\mu$-strong convexity the exponent halves, to $\varepsilon^{-\gamma/2}$, because the iterates settle at a noise floor and the bias budget relaxes accordingly. Both bounds are independent of the step size, and hence of the smoothness constant, and we show that the cost is a functional of the underlying gradient flow rather than of any discretization of it.
Siqing Liu, Eric Chung, Yiran Wang
Thu 24 Sep
math.NAcs.NAstat.ML
pdf html
- favorites near “Monte Carlo Geometry Processing” 0.83
- reading gp, online, basis, gives
- interest Kernels, operators and functional data 0.77
Abstract
Residual-based adaptive GMsFEM for high-contrast elliptic problems repeatedly evaluates local weighted $H^{-1}$ indicators on every coarse neighborhood, making indicator evaluation a recurring cost in repeated-query settings. We introduce a non-intrusive Gaussian-process (GP) surrogate for the indicator scores used in D\"orfler marking. The operational predictor is the GP posterior mean, algebraically equivalent to a kernel ridge regression (KRR) estimator under the stated convention; it uses compressed local solution and spectral features without changing the multiscale solve, local spectral construction, or basis enrichment. A nonuniform perturbed-marking result quantifies how pointwise score errors affect the exact indicator mass captured by surrogate-selected neighborhoods, while a conditional bounded-discrepancy KRR pathway identifies sufficient assumptions for such score bounds. In controlled held-out in-distribution tests, the surrogate-guided method gives error-versus-DoF trends comparable with classical $H^{-1}$-residual offline adaptivity and evaluates the online indicator component 2.0-2.1 times faster, excluding offline data generation and GP training.
Jing Qin
Mon 28 Sep
stat.MEstat.AP
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.84
- interest Semiparametric inference and debiased machine learning 0.84
- favorites near “Orthogonal Statistical Learning” 0.82
Abstract
The standard dependence summaries used in biomarker studies -- Pearson's r, Spearman's rho, Kendall's tau -- take values in [-1, 1] with 0 indicating no linear or monotone association. Zero does not distinguish independence from non-monotone dependence, so the scale cannot represent threshold effects, heteroscedasticity, and tail shifts common in biomarker practice. We formulate independence testing as a composite Bernoulli likelihood ratio: at each threshold t, comparing the Bernoulli laws of 1(Y <= t) conditionally on X versus marginally, aggregated over cut points. The resulting coefficient xi_cut lies on [0, 1] with 0 iff X and Y are independent (under continuity of Y) and 1 iff Y is a measurable function of X. Fisher weighting arises at second order from the Bernoulli likelihood, and xi_cut equals twice the threshold-averaged mutual information between X and 1(Y <= t), giving a distribution-free lower bound on I(X; Y). A second-order expansion recovers the Fisher-weighted Dette-Siburg-Stoimenov measure, which coincides under continuity with Chatterjee's rank correlation. Estimation uses a Nadaraya-Watson plug-in with a max-over-grid bandwidth; inference is by exact permutation. In biomarker-motivated simulations T_cut substantially outperforms rank-based coefficients on W-shaped non-monotone and heteroscedastic alternatives. We illustrate on the Seattle cohort (n=70, ages 21-88) of the aging plasma proteome dataset, screening all 1,305 proteins for age dependence: under Benjamini-Hochberg control at q<0.05, T_cut rejects on 70 proteins, six of which are missed by Pearson, Spearman, and Chatterjee at the same FDR level.
Luyang Fang, Haoran Lu, Jiazhang Cai, Tao Wang, Huimin Cheng, Wenxuan Zhong, et al.
Tue 29 Sep
stat.MLcs.LG
pdf html
- reading theory, language model, conceptual, viewed
- interest Language models as statistical objects 0.80
- your work near “Valid Inference with Imperfect Synthetic Data” 0.77
Abstract
Knowledge Distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient ``student'' counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive settings. Although KD is widely used in practice, it is often viewed primarily as an engineering technique, with a unified statistical perspective remaining less developed. This review bridges that gap by presenting a unified Bayesian formulation of KD that formulates teacher predictions as prior information. This provides a principled interpretation of how teacher information is incorporated into student learning and establishes a rigorous connection to uncertainty quantification. We demonstrate how this foundational lens reconciles classical distillation with modern extensions in generative and foundation-model systems, showing that contemporary developments remain rooted in these same statistical principles. By synthesizing theory with emerging methodologies and diverse applications, this review provides a conceptual roadmap and identifies critical open problems for the future of the field.
Swapnaneel Bhattacharyya, Aaditya Ramdas
Thu 24 Sep
math.STstat.MEstat.MLstat.TH
pdf html
- interest Adaptive data collection and anytime-valid inference 0.83
- reading martingales, theory, control, sequential
- your work near “Timing as an Action: Learning When to Observe and Act” 0.76
Abstract
We study distribution-free sequential changepoint detection for independent observations with unknown and unrestricted pre- and post-change laws. We build on the conformal test martingales and associated e-detectors of Vovk(2021), which control the probability of false alarm (PFA) and the average run length (ARL) respectively. The majority of these works focus on validity, with statistical efficiency usually left for simulations. We develop a comprehensive theory of how conformal p-values behave under non-exchangeable data with a changepoint at an unknown time $T$. We use this to analyze the post-change growth and resulting detection delay of conformal martingale methods, and prove that the standard existing methods are suboptimal for PFA and ARL control, and can lead to delays that are $\Omega(T)$ and $\Omega(\sqrt{\text{ARL}})$ respectively. We propose different conformal e-processes and e-detectors that are provably minimax optimal, with delays $\Theta(\log T)$ and $\Theta(\log \text{ARL})$ respectively, and have much shorter delays in simulations.
Alessandro Trenta, Riccardo Massidda, Davide Bacciu, Sara Magliacane
Wed 30 Sep
cs.LGcs.AIstat.ML
pdf html
- reading equations, odes, equation, causal
- favorites near “Spectral State Space Models” 0.76
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.76
Abstract
We study the problem of recovering the governing ODE of a dynamical system from unstructured, high-dimensional observations such as images. Existing methods for ODE discovery typically assume direct measurements of the variables, or do not provide theoretical guarantees on the learned variables and equations. While Causal Representation Learning (CRL) methods provide guarantees on identifying variables from high-dimensional observations up to component-wise diffeomorphisms, we show that in general these variables cannot be used directly as input to equation discovery methods, which typically assume that the variables will lead to sparse equations. So we introduce SParse Equivalent Equation Discovery AutoEncoder (SPEED-AE), a framework that combines a pretrained CRL method with a component-wise autoencoder that learns transformations of variables that are amenable to sparse ODE discovery. We show that for polynomial ODEs, this additional step allows us to restrict the identifiability of each variable from polynomial to monomial diffeomorphisms. Experiments on Lotka-Volterra, Lorenz, and a two-pendulum system show that SPEED-AE improves on the disentanglement of the CRL methods and that it recovers ODEs that are closest to the ground truth, while achieving state-of-the-art forecasting performance.
Sakshi Arya, Cheng Soon Ong
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Preference-based Online Learning with Dueling Bandits: A Survey” 0.84
- interest Reinforcement learning, policy learning and off-policy evaluation 0.80
- your work near “Timing as an Action: Learning When to Observe and Act” 0.75
Abstract
We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (NBL), a greedy procedure that uses a sequential Stein contrast to learn the optimal boundary directly, without estimating the reward functions or the common link. We characterize the local Riemannian dynamics of NBL through a decision stability coefficient balancing arm separation, link geometry, and the context distribution. We show that this stability is connected to the elicitation geometry of the underlying convex potential. Under local decision stability, NBL contracts toward the optimal boundary and achieves $O(\log n)$ expected regret. Numerical experiments illustrate the predicted stability regimes and compare NBL with a parametric greedy benchmark under link misspecification.
Shoichi Otomo
Tue 29 Sep
cs.CYecon.EM
pdf html
- reading economic, macroeconomic, economy, economies
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.76
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.65
Abstract
This manuscript is an English translation and extended version of a paper originally published in Japanese (2022). Driven by remarkable advances in remote sensing and big data processing, spatial technologies are increasingly leveraged in economic research. While satellite nighttime light intensity is widely recognized for tracking macroeconomic parameters - such as regional GDP, employment, and population - less attention has been paid to fine-grained spatial processing methodologies for local tourism economies. This study first details a raster processing technique applied to nightlight imagery. It then focuses on Yuzawa Town (Uonuma District, Niigata Prefecture), analyzing the local economy and sports/tourist attractions. Specifically, I investigate the spatio-temporal relationships between tourist arrivals and various local statistical datasets. Furthermore, this paper highlights how nightlight data can capture intra-municipal economic dynamics that remain undetectable through standard macroeconomic indicators.
Rajit Rajpal, Benedict Leimkuhler
Mon 28 Sep
stat.MLcs.LG
pdf html
- reading friction, damping, continuous time, momentum
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.80
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
Abstract
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as Adam's second-moment buffer. Here we replace iKFAD's friction tensor $\xi$ with a rank-1 outer-product factorisation built from row and column momentum statistics, resulting in Rank-1 iKFAD (R-iKFAD). This reduces the friction memory footprint from $\mathcal{O}(mn)$ to $\mathcal{O}(m+n)$ per layer, which approximately halves iKFAD's total optimiser state. Despite this reduction, R-iKFAD maintains parity in performance with iKFAD: experiments on GPT2-Nano, TinyViT, DistilBERT and GPT2-S confirm that it matches or exceeds iKFAD while nearly halving the memory footprint and remaining comparably robust to hyperparameters. We analyse the continuous-time dynamics in two damping regimes. For linear damping ($\gamma>0$) we prove exponential convergence under strong convexity. For $\gamma=0$, the preferred option in our experiments, the friction is generated entirely from past momentum and switches off as the momentum vanishes, so geometric convergence cannot be shown. We nonetheless prove convergence to the minimiser, together with matching upper and lower bounds on the energy: of order $t^{-1}$ when the regularisation scale $\epsilon_{\mathrm{stab}}$ is zero, and of order $t^{-1/2}$ when it is positive. To our knowledge this is the first convergence rate for a rank-1 factored optimiser in continuous time, and the first such result that does not require positive damping.
Yaiza Bermudez, Samir M. Perlaza, Iñaki Esnaola
Fri 25 Sep
stat.MLcs.LG
pdf html
- reading empirical, data points, subject, special
- interest Online learning and statistical learning theory 0.83
- favorites near “Optimal Decision Rules for Weak GMM” 0.84
Abstract
In this paper, a method for achieving exact unlearning for Gibbs supervised learning algorithms is proposed using a variational formulation inspired by empirical risk minimization subject to relative entropy regularization (ERM-RER). Such a method consists of maximizing the expected empirical risk over the dataset to be unlearned subject to a regularization by relative entropy with respect to the original algorithm. The optimization variable is a probability measure on the models; and the solution is another Gibbs probability measure that represents a new Gibbs supervised learning algorithm. The method guarantees exact unlearning in the sense that the new Gibbs algorithm coincides in distribution with the algorithm that would have been obtained by retraining from scratch on the dataset to be retained. As a byproduct, a framework for reweighting data points in ERM-RER by strategically choosing both the reference measure and the regularization factor is obtained. In this framework, exact unlearning is the special case in which zero-weight is assigned to the contribution of the data points to be unlearned. More generally, depending on the choice of certain parameters, data points can be up-weighted or down-weighted in ERM-RER problems for particular purposes, e.g., controlling the generalization error of Gibbs algorithms. This paves the way for new constructive or adversarial views on classical reweighting data points in ERM-RER.
Haris Moazam Sheikh
Tue 29 Sep
cs.LGmath.OCstat.ML
pdf html
- favorites near “Monte Carlo Geometry Processing” 0.80
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- interest Kernels, operators and functional data 0.79
Abstract
Multi-objective Bayesian optimization (MOBO) is commonly approached through specialized acquisition functions or scalarization schemes designed to explicitly account for trade-offs among non-preferential objectives. In this work, we show that such complexity might be unnecessary. We propose a framework that extends standard single-objective acquisition functions directly to the multi-objective setting through a hypervolume-based transformation. We further extend hedge strategies for acquisition functions, which are typically used only in single-objective optimization, to the multi-objective regime. Our approach requires minimal modification to existing Bayesian optimization pipelines and avoids the need for bespoke multi-objective formulations. We demonstrate how a broad class of commonly used single-objective acquisition functions and hedge strategies can be adapted in a principled manner to handle multiple objectives, while preserving their intuitive interpretation and computational efficiency. Empirically, we evaluate the proposed methods across a range of synthetic and real-world multi-objective benchmarks. Despite their simplicity, our extensions consistently match or outperform more complex state-of-the-art MOBO methods in terms of optimization performance and sample efficiency. These results suggest that effective multi-objective Bayesian optimization can be achieved by reusing and carefully extending well-established single-objective acquisition strategies, offering a simpler and more flexible alternative to existing approaches.
Miguel Ángel Beltrán-Sánchez, Miguel Ángel Martínez-Beneito, Terje Eikemo, Sara Martino, Andrea Riebler
Fri 25 Sep
stat.AP
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.75
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.73
- reading social, good, measured, straightforward
Abstract
Survey responses are often measured using ordered response categories. In the European Social Survey, self-rated health is measured on a five-point scale from very good to very bad, yet analyses commonly dichotomise responses into binary categories of "good" and "poor" health. Binary indicators provide prevalence measures that are straightforward to communicate, but dichotomisation reduces information and may limit captured health variation, while the cut-off may influence estimates and substantive conclusions. Systematic evidence on these effects in spatial settings remains limited. We address this gap using European Social Survey round 11 (2023/2024) for Western Europe. We model self-rated health among male respondents by age, education and region. Bayesian spatial individual-level models with poststratification provide population-representative estimates. We compare an ordinal cumulative logit model for the five-category outcome with Bernoulli logistic regression models using two dichotomisations differing in the classification of "fair" health. Older age and lower education are consistently associated with worse self-rated health across specifications, suggesting relatively robust fundamental age and educational gradients. However, their magnitude and uncertainty, and some geographical conclusions, are sensitive to how the outcome is modelled. Assigning "fair" health to either side of a binary cut-off changes the regions identified as having above-average levels of less favourable health. The ordinal model retains category-specific information and can produce familiar binary prevalence estimates through aggregation. Binary indicators remain useful, particularly for monitoring and communication. Nevertheless, the selected cut-off should be justified and sensitivity to alternative cut-offs or an ordinal modelling approach should be considered, especially for geographical comparisons.
Wei Cao, Shanshan Wang
Tue 29 Sep
stat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.82
- interest Adaptive data collection and anytime-valid inference 0.82
- reading online, procedures, heterogeneity, sequential
Abstract
Streaming data, characterized by high volume, rapid arrival rates, and evolving distributions, have become increasingly prevalent in modern applications. Developing efficient and reliable estimation procedures is therefore essential for real-time statistical analysis. However, most existing online estimation methods rely on the assumption of batch homogeneity, which can be violated in practice due to abnormal batches, distributional shifts, or other forms of batch heterogeneity. To address this challenge, we develop renewable online expectile regression procedures for heterogeneous streaming data. Specifically, we propose two complementary strategies for handling abnormal batches: (1) a detection-based approach that employs a sequential monitoring mechanism based on score test statistics to identify and remove potentially abnormal batches; and (2) an adaptive-weighting approach that assigns data-driven weights to incoming batches, reducing the influence of abnormal or drifting batches while retaining information from reliable observations. Both strategies rely solely on score test statistics and can be seamlessly integrated into existing renewable estimation and inference frameworks without requiring additional structural assumptions. Furthermore, to enhance robustness against heavy-tailed errors and outliers, we replace the conventional l2 loss with the Huber loss and develop a robust extension of renew?able online expectile regression. Extensive simulation studies and analyses of clinical datasets demonstrate that the proposed methods achieve improved estimation accuracy and robustness in the presence of batch heterogeneity. Overall, the proposed framework provides a flexible and effective solution for renewable expectile regression in complex streaming data environments.
Jie Deng
Fri 25 Sep
cs.LGstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.81
- interest Adaptive data collection and anytime-valid inference 0.80
- favorites near “Data Science at the Singularity” 0.82
Abstract
A released table is often treated as an i.i.d. sample, although its repeated rows may encode business frequency, repeated entities, joins, resampling, or extraction errors. We show that this ambiguity creates a hidden measurement layer with three consequences: feature-identical rows impose an attained evaluation ceiling, row-weighted AUROC is sensitive to replication, and row-trained detectors learn a multiplicity-size-biased law. An exact-row audit of all 690 OddBench datasets finds train-test overlap in 355, feature-identical label conflict in 147, and a test anomaly identical to a training normal in 137. Switching from row to support weighting changes AUROC by at least 0.05 on 50-61 datasets across four classical detector geometries. We introduce SCOUT (Support-Count Orthogonalized Unsupervised Testing), a factorized anomaly detector that separates replication-invariant support evidence from exposure-aware count evidence. Factorwise split-conformal calibration yields marginal false-positive-rate control, while the support channel is exactly invariant to arbitrary positive row replication. On 686 OddBench datasets and five seeds, support-only SCOUT is non-inferior to row-wise Isolation Forest in raw AUROC and improves replication-invariant AUROC. External normal-support evaluations track nominal false-positive levels, and four backbones remain exactly unchanged under controlled replication. Semi-synthetic interventions show that conditional count modeling helps materially only under strong rate heterogeneity. These results specify when multiplicity should be treated as signal, nuisance, or uninterpretable without additional information.
Masoud Kavian, Abdellatif Zaidi, Milad Sefidgaran
Tue 29 Sep
stat.MLcs.ITcs.LGmath.IT
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.81
- interest Language models as statistical objects 0.82
- your work near “Timing as an Action: Learning When to Observe and Act” 0.75
Abstract
We develop an information-theoretic framework for generalization in next-token prediction under temporally dependent data. We consider independent trajectories generated by finite-memory Markov processes and distinguish algorithmic dependence, quantified by mutual information, from temporal dependence, characterized by mixing. For cross-entropy loss, we derive an expected generalization bound using the Donsker--Varadhan variational representation and a McDiarmid-type concentration inequality for Markov chains. A refinement captures the joint effect of context length and temporal mixing through the mixing properties of the history-state process. We then extend the bound through a rate--distortion formulation, replacing mutual information with the minimum information rate required to represent the learned model within a prescribed distortion in the generalization gap, yielding informative guarantees for deterministic algorithms over continuous hypothesis spaces. For margin-based prediction, we derive explicit bounds for linear and self-attention next-token predictors via noisy low-dimensional compression, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing. Experiments on TinyStories and ETTh2 show that longer contexts can reduce both training and test losses, but typically reduce training loss more, enlarging the generalization gap. A complementary ETTh2 analysis identifies an effective predictive-memory scale near 24 hours, with no statistically supported improvement beyond this scale, offering a plausible explanation for test-performance saturation at larger contexts.
Tanabodee Monggon, Songkiat Sumetkijakan, Tippawan Santiwipanont, Monchai Kooakachai
Fri 25 Sep
stat.ME
pdf html
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.75
- reading computations, main, complex, issues
- interest Causal inference and identification in econometrics 0.79
Abstract
Traditional multivariate copulas often fail to capture complex dependence structures and may impose restrictive assumptions that limit their applicability to real-world data. To address these issues, vine copulas provide a flexible framework by constructing multivariate models from bivariate copulas. The three main types are D-vines, C-vines and R-vines. To simplify computations, the simplifying assumption is often applied, assuming conditional copulas are independent of conditioning variables. This results in some variables exerting minimal influence on the conditional copulas. Therefore, this study explores two subclasses of R-vines with node degrees capped at three, assessing its performance against existing models. The statistical results demonstrate that the proposed degree-constrained vine structures can serve as effective alternatives to existing vine copula models.
Hui Yang, Yun Chen, Fabio Leonelli
Tue 29 Sep
eess.SYcs.SYstat.AP
pdf html
- favorites near “Spectral State Space Models” 0.78
- reading nonlinear, attempts, rise, developments
- interest Kernels, operators and functional data 0.78
Abstract
Nonlinear dynamics arise whenever multifarious entities of a system cooperate, compete, or interfere. For example, cardiovascular system involves a great level of complexity. Multi-lead ECG signals are generated through orchestrated depolarization and repolarization of cells and manifest significant nonlinear dynamics. Nonlinear dynamical systems defy understanding based on the traditional reductionist's approach, in which one attempts to understand a system's behavior by combining all constituent parts that have been analyzed separately. In order to cope with system complexity, modern healthcare systems are investing in advanced physiological sensing and patient monitoring, thereby giving rise to big data. Realizing the full potential of big data for healthcare intelligence requires fundamentally new methodologies to harness and exploit complexity. However, available nonlinear dynamics techniques are either not concerned with healthcare analytical objectives or fail to effectively analyze big data to extract useful information for improving healthcare services. There is an urgent need to develop analytical methodologies that fully exploit the underlying nonlinear dynamics in physiological systems for advancing healthcare services with exceptional features such as personalization, responsiveness, and superior quality. This chapter presents some theoretical developments and tools to advance the applications of nonlinear dynamics principles in health care. Specifically, we focus on sensor-based characterization and modeling of nonlinear dynamics (i.e., multifractal analysis and multiscale recurrence quantification). Then, current developments and applications of these methodologies are examined for characterizing and exploiting heart rate variability and space-time ECG signals.
Abbas Mammadov, Jerry Y. Huang, Justin Lin, Partha Kaushik, Sheel Shah, Kartik Nair, et al.
Thu 24 Sep
cs.LGcs.CVstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.79
- interest Language models as statistical objects 0.80
- reading higher, reinforcement learning, reinforcement, recipe
Abstract
Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. We show that the resulting problem is equivalent to a deterministic optimal control problem on the flow. Given a pre-trained flow map, this equivalence yields a simulation-free reinforcement learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output is a fine-tuned flow map that retains strong reward-aligned performance at few-step inference budgets without post-hoc distillation. Experiments on ImageNet-256 and text-to-image show that WTF achieves higher reward with comparable or higher diversity than baselines, while requiring up to $280\times$ less training compute. More broadly, we argue that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation is only one of many choices worth revisiting.
Xiaojing Sun, Yuhan Zeng, Zihua She, Xiao Wang
Tue 29 Sep
stat.MLcs.AIcs.LGstat.ME
pdf html
- reading theory, control, empirical, advances
- your work near “Efficient Online Estimation of Causal Effects by Deciding What to Observe” 0.79
- interest Adaptive data collection and anytime-valid inference 0.81
Abstract
As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite resources are repeatedly reused, new candidates are proposed based on feedback from the same evaluation set, so the search trajectory can adaptively overfit and empirical improvement may not reflect genuine population improvement on the underlying task distribution. Some existing methods account for multiple comparisons but assume that candidates are chosen independently of the evaluation set, and therefore do not control this adaptive dependence. To address this, we propose REUSE (Risk-controlled Evaluation Under Sequential Evolution), a certified evaluation and promotion framework that allows a fixed evaluation set to support repeated adaptive decisions while providing statistical guarantees. For a user-specified error level $\alpha$, with probability at least $1-\alpha$, every promoted modification is a genuine population improvement on the underlying task distribution. REUSE achieves this by strictly limiting the evaluation feedback returned to the search process and accounting for possible promotion histories within the error budget. We develop detailed statistical theory for RSI evaluation in this setting, including simultaneous error control, valid lower bounds on cumulative improvement, and a characterization of the fundamental limits of adaptive evaluation reuse. In live self-improvement experiments, REUSE commits substantially fewer false promotions than evaluation frameworks from current RSI systems and error-controlled baselines, reducing the proportion of false promotions from up to 20.7% to 0%, while achieving final true population performance comparable to the best baselines.
Linard Hoessly
Fri 25 Sep
stat.AP
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.81
- reading perfect, outcome, true, approaching
- interest Reinforcement learning, policy learning and off-policy evaluation 0.78
Abstract
At a fixed threshold, relative utility (RU) measures the net-benefit gain of a prediction model over the better of treat-all and treat-none relative to the corresponding gain under perfect outcome classification. We illustrate that RU equal to 1 therefore represents perfect outcome classification at that fixed threshold, not perfect probabilistic prediction. However, even when every predicted probability equals the true probability, observed RU can equal 0. In a simple constant-risk setting, this occurs with probability approaching 1 as the sample size increases. Consequently, the distance from observed RU to 1 should not in general be interpreted as improvement achievable by a better prediction for binary probabilities.
Soham Dan
Tue 29 Sep
stat.MLcs.LG
pdf html
- reading interpolation, theory, argument, come
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.77
- favorites near “Random Walks on Simplicial Complexes and Harmonics” 0.82
Abstract
Testing whether two populations of networks share the same edge probabilities is a basic problem in network inference. How hard it is depends on the norm used to measure the difference. For the inhomogeneous Erd\H{o}s--R\'enyi (IER) model, the optimal sample complexity is known for every integer $L_r$ norm and for $1\le r<2$. For non-integral $r>2$, however, the known upper and lower bounds do not match, and the lower bound was conjectured to be tight. We study this gap for two-sample testing on aligned vertices. We propose a test that runs two published statistics, of orders $2$ and $\lceil r\rceil$, on the same data and rejects if either one rejects. Its thresholds come from H\"older interpolation, so that both statistics have the same sample cost. We prove that this test attains the conjectured rate. Combined with earlier results, this shows that for every fixed $r\ge1$ the minimax sample complexity is of order $n^{\max\{4/r-1,\,2/r\}}/\epsilon^2$, even when the separation changes with $n$. In simulations with $n$ between 32 and 256, the number of graphs needed for 80\% power at level $0.05$ grows with $n$ at a rate consistent with the theory. For $r=2.5$, for example, the fitted exponent is $0.78$, against the theoretical value $0.8$. Interestingly, the two statistics split the work as the interpolation argument suggests: the higher-order statistic is more powerful when only a few edges change, and the $L_2$ statistic when many edges change.
Etienne Boursier, Nicolas Flammarion
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.83
- interest Kernels, operators and functional data 0.79
- reading target parameters, empirical, modern, implicit
Abstract
Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has $m+1$ features, $m$ of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only $\mathcal{O}(d)$ fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.
Junxuan Li, Arko Mukherjee, Soumyabrata Pal
Tue 29 Sep
cs.AIstat.AP
pdf html
- reading wrong, summarization, rates, humans
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.79
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.80
Abstract
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Takashi Furuya, Nicholas H. Nelsen, Frank Cole
Tue 29 Sep
cs.LGcs.NAmath.NAstat.ML
pdf html
- reading operators, operator, uniformly, continuous
- interest Kernels, operators and functional data 0.78
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.69
Abstract
Many learning tasks map an input distribution to an output distribution. A natural way to model such an operator is to transform each input sample using a continuous function that may depend on the entire input distribution, and then take the distribution of the transformed samples. This defines a measure-dependent pushforward model and includes measure-theoretic formulations of transformers. We ask when such models can approximate arbitrary continuous operators between spaces of probability measures. We first show that universal approximation fails when atomic inputs are allowed: some continuous measure-to-measure operators that split or redistribute atomic mass cannot be approximated arbitrarily well by deterministic pushforward models. We then introduce the uniform level set condition, which requires a continuous measure-dependent scalarization whose shrinking level set neighborhoods carry uniformly vanishing mass over the input family. This condition is satisfied, in particular, by compact families of absolutely continuous measures. On every compact family satisfying this condition, we prove that any continuous measure-to-measure operator with outputs of finite $p$-th moment can be uniformly approximated, in the $p$-Wasserstein distance, by continuous measure-dependent pushforwards. Combining our theorem with existing approximation results for measure-dependent in-context maps yields universal approximation by measure-theoretic transformers. We also extend the framework to continuously-varying source measures, yielding a corresponding universality result for a class of pushforward models that are closely aligned with cross-attention architectures.
Pierdomenico Duttilo, Francesco Lisi
Fri 25 Sep
stat.AP
pdf html
- reading economic, semiparametric, forecasting, functional
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.73
- interest Macroeconometrics and time series 0.76
Abstract
Accurate tide forecasts are essential for coastal management, navigation, flood-risk reduction, and infrastructure protection. Observed sea level can be decomposed into astronomical and non-astronomical components, the latter mainly driven by meteorological effects. This study investigates a hybrid framework for hourly sea-level forecasting that combines harmonic analysis (HA) for the astronomical component with data-driven models for the non-astronomical contribution. The approach is evaluated at six tide-gauge stations with different tidal regimes: Venice, Trieste, Saint-Malo, Vard{\o}, Nikiski, and Nagasaki. Four data-driven model classes are considered: (i) linear parametric models, represented by autoregressive models with exogenous variables; (ii) functional parametric models, based on functional autoregressive models with exogenous variables; (iii) semiparametric and nonlinear models, including generalized additive models and autoregressive neural networks; and (iv) a semi-functional non-standard k-nearest-neighbours approach combining similarity in recent non-astronomical trajectories and meteorological conditions. Results reveal that hybrid models reduce forecast errors by 52.9-54.9% on average relative to HA. The generalized additive model is the most competitive across locations, while k-nearest neighbours performs best at Saint-Malo and the autoregressive model with exogenous variables is favoured in Nagasaki. For Venice, an economic decision-making case study assesses the operational use of sea-level forecasts in managing the MoSE flood-barrier system.
Alexander Henzi, Tiange Liu, Xinwei Shen
Fri 25 Sep
stat.MEcs.LGstat.ML
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.78
- favorites near “Spectral State Space Models” 0.77
- interest Language models as statistical objects 0.80
Abstract
We propose Sufficiently Reduced Distributional Regression (SRDR), a generative method that combines conditional distribution estimation with nonlinear sufficient dimension reduction (SDR). It builds on a characterization of sufficiency through strictly proper scoring rules: a dimension reduction is sufficient if and only if predicting the response from the reduced covariates incurs no loss in expected score relative to the full covariates. Sufficient dimension reduction thus becomes a risk minimization problem. SRDR jointly trains a dimension reduction map and a generative prediction model by minimizing the energy score, which can be estimated by sampling without density evaluation or adversarial training. The framework extends to multi-environment data and to classification. We prove that the estimated conditional distributions converge in energy distance to the true ones, which implies that the learned representation is asymptotically sufficient. In simulations and applications to CT slice localization, superconductivity, and digit classification, SRDR recovers low-dimensional sufficient structure and matches or outperforms state-of-the-art nonlinear SDR methods in representation quality and predictive performance.
Yulong Li, Ke Wan, Toshio Shimokawa, Kensuke Tanioka
Wed 30 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.84
- interest Empirical Bayes, shrinkage and decision theory 0.81
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.81
Abstract
Interpretable prediction models are important in biomedical research, where predictive accuracy must often be balanced against the ability to examine predictor-outcome relationships. Automatic Binary Logistic Estimation (ABLE) provides interpretable rule-based representations by constructing additive models using threshold rules. However, the original method is restricted to binary outcomes, and its greedy forward construction follows a search path, making candidate rules dependent on earlier selections and potentially excluding alternative structures. We extended ABLE to binary, continuous, and time-to-event outcomes using outcome-specific fitting criteria, yielding Generalized ABLE (GABLE). We then propose bootstrap rule pooling for GABLE (BRP-GABLE), which applies GABLE rule generation to bootstrap samples to explore alternative search paths, pool distinct rules, and perform a single global LASSO selection over the pooled rules and truncated linear terms. Unlike conventional bagging, bootstrap-specific models are not averaged, and resampling is used to diversify the candidate rule space before constructing a final model. Simulations showed that BRP-GABLE consistently improved predictive performance compared with GABLE and provided competitive prediction with relatively simple rule structures. Application to overall survival demonstrated its use for clinically interpretable prediction. BRP-GABLE provides a flexible framework for constructing sparse and interpretable predictive models across multiple outcome types.
Srijato Bhattacharyya, Huiyan Sang, Bani Mallick
Tue 29 Sep
stat.ME
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.79
- favorites near “Monte Carlo Geometry Processing” 0.78
- interest Empirical Bayes, shrinkage and decision theory 0.80
Abstract
Clustering problems increasingly involve complex objects observed over space, such as distributions, matrices, functions, images, or multivariate data, for which a scientifically meaningful dissimilarity between objects is often easier to specify and computationally more tractable than an object-response likelihood model. We propose DISCCO, a Bayesian framework for clustering spatially indexed complex objects using only a pairwise distance matrix and a spatial adjacency graph, with broad applicability and minimal user modeling requirements. The model combines a hierarchical distance-based likelihood accounting for within-cluster compactness and between-cluster separation with a random spatial graph partition prior, ensuring that posterior clusters are spatially contiguous. A key model feature is a set of node-specific frailty parameters that induce dependence among overlapping within-cluster distances and provide posterior summaries of object-level centrality or peripherality within each inferred cluster. We develop a partially collapsed Markov chain Monte Carlo algorithm for posterior inference. Simulations with distribution- and matrix-valued responses show that the proposed spatial distance-clustering framework improves region recovery relative to existing distance-clustering methods, while the frailty layer provides interpretable centrality summaries. Real applications to Houston Census Block Group racial-composition distributions and Western US county-level cancer mortality matrices illustrate how the method recovers interpretable contiguous clusters and frailty-based centrality maps.
Prasanjit Dubey, Aritra Guha, Xiaoming Huo
Tue 29 Sep
stat.MLcs.LG
pdf html
- reading questions, documents, little, nodes
- your work near “Valid Inference with Imperfect Synthetic Data” 0.77
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.78
Abstract
Retrieval-augmented generation (RAG) lets language models answer questions more accurately by consulting relevant documents. Many valuable collections, such as medical records, cannot be pooled because of privacy rules. Federated RAG leaves each collection with its owner, or node, which scores candidate answers from its own documents; a central hub combines the scores. Some nodes, called Byzantine, may be compromised, faulty, or misled by instructions hidden in documents, and report arbitrary scores. Conformal prediction returns a set containing the correct answer with a chosen probability, using a cutoff set in a calibration step on questions with known answers. An unknown group of nodes, no larger than a declared bound, may misreport both in this step and at query time. Existing methods assume every node is honest or protect only the calibration step. We observe that the honest nodes are the same in both steps. The hub therefore has all nodes score the same calibration questions, and keeps a candidate only if some plausible group of honest nodes, using its own scores in both steps, would keep it. We prove that the resulting sets contain the correct answer with the chosen probability in finite samples, whatever the Byzantine nodes report. No method using the same information can return smaller sets without risking the loss of an answer the honest nodes support. If nodes fail at random, the guarantee weakens only by the probability that more nodes fail than declared. In simulations, on real question-answering tasks including medical exams, and with language models as nodes, some hijacked, our sets reached the target whenever no more nodes misbehaved than declared, while plain averaging could miss it. They were also clearly smaller than those of simpler methods with the same protection, most of all when the declared bound was generous, so a cautious bound costs little.
Joachim Freyberger, Julius Kappenberg
Tue 29 Sep
econ.EM
pdf html
- reading economic, inequality, choice, inequalities
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- interest Bayesian computation and structural estimation 0.78
Abstract
Economic theory frequently implies linear inequality restrictions on parameters or functions of interest. A common way to impose such restrictions is to project an unrestricted estimator onto the feasible set. Projection estimators arise naturally from constrained least squares, instrumental variables, generalized method of moments, maximum likelihood, and related extremum procedures. When the sampling covariance, loss function, and projection criterion induce different geometries, the choice of projection geometry can substantially affect risk. We study this choice in a fixed-dimensional local Gaussian experiment under quadratic loss. At exact-boundary configurations where only one maintained inequality binds, inverse-covariance projection is pointwise optimal. When at most two inequalities are locally relevant, it weakly improves on the unrestricted estimator throughout the corresponding local experiment and is minimax over exact-boundary configurations. For an arbitrary number of inequalities, we provide a sufficient condition for boundary minimaxity, but show by counterexample that inverse-covariance projection need not be boundary minimax once three inequalities can bind. Motivated by these results, we propose selecting the projection geometry to minimize worst-case boundary risk subject to a local no-harm condition relative to unrestricted estimation. We develop a feasible implementation and study its finite-sample performance in simulations and an application to gasoline demand.
Cristina Galán-Arcicollar, Danilo Alvares, Josu Najera-Zuloaga, Dae-Jin Lee
Fri 25 Sep
stat.AP
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.82
- reading submodel, processes, bayesian joint, dynamic
- interest Causal inference and identification in econometrics 0.77
Abstract
Questionnaire-based patient-reported outcomes (PROs) are discrete, bounded and overdispersed, yet joint models relating them to survival may ignore these features or estimate both processes sequentially. We propose a Bayesian joint model combining a beta-binomial mixed-effects submodel with a Weibull proportional hazards submodel, linked through the subject-specific response probability. Simulations show that simultaneous estimation reduces bias in the longitudinal slope and yields practically unbiased association estimates, unlike two-stage estimation. In a cohort of 543 patients with chronic obstructive pulmonary disease, the model identified associations for all eight SF-36 dimensions and for two of three SGRQ dimensions, including several associations not detected by the two-stage approach, and provided dynamic survival predictions.
Jifan Li, Ning Ning
Tue 29 Sep
stat.MEcs.CVcs.LG
pdf html
- interest Macroeconometrics and time series 0.81
- favorites near “Monte Carlo Geometry Processing” 0.80
- reading forecasting, gains, future, temporal dependence
Abstract
Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architecture supporting trend, seasonal and cycle dynamics. Visual compression reduces representation dimension, while Bayesian variable selection reduces active regression dimension. Posterior prediction combines forecasts across predictor subsets using their posterior probabilities as weights and accounts for parameter uncertainty and future disturbances. The model forecasts joint visual--physical states recursively and physical targets directly. Across four forecasting tasks spanning object motion, vegetation greenness and solar power, ViBR-WM achieves lower mean overall physical-target error than Temporal Straightening, ConvLSTM, PredRNN and SimVP on every task. Repeated fitting and resampling support these overall gains.
Xiaodong Li, Zhentao Li
Thu 24 Sep
stat.MLcs.LG
pdf html
- interest Semiparametric inference and debiased machine learning 0.84
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- favorites near “Operator-valued Kernels for Learning from Functional Response Data” 0.82
Abstract
We study multitask regression when coefficient sharing can differ by predictor. For a given predictor, many tasks may have the same coefficient while a few differ, and the exceptional tasks need not be the same for another predictor. We describe this structure by two quantities: the number of active predictors and the total number of task coefficients that differ from the most common value for their predictor. We estimate the coefficient matrix by penalizing all pairwise coefficient differences across tasks, with an additional group penalty when predictor selection is needed. The resulting upper and lower bounds have the same dependence on these two quantities. We also consider the stronger setting in which a large set of tasks shares one entire coefficient vector. Under explicit sample-size conditions, the same pairwise estimator pools those tasks exactly, while allowing the remaining tasks to differ. Simulations and household energy data illustrate the transition between broad sharing and task-specific coefficients.
Minwoo Jang, Jaechang Kim, Minhyeon Oh, Jeongyeon Hwang, Jungseul Ok
Tue 29 Sep
cs.AIcs.CLcs.CRcs.LGstat.ML
pdf html
- reading questions, reinforcement learning, distilled, preference optimization
- your work near “Valid Inference with Imperfect Synthetic Data” 0.78
- interest Language models as statistical objects 0.77
Abstract
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However, a distilled model may undergo further SFT, preference optimization, or reinforcement learning before release, while an auditor may lack access to the pre-distillation checkpoint required by reference-based attribution. To close this gap, we propose SCOUT, an output-only method that aggregates recurring *syntactic patterns* into candidate profiles, filters low-contrast patterns, and calibrates student--candidate distances against inter-candidate distances. SCOUT supports attribution and abstention using only current texts, without model weights, token likelihoods, or historical checkpoints. Auditing publicly released descendants of distilled models spanning diverse post-training objectives, SCOUT consistently identifies the distillation source. Furthermore, tracing teacher-associated *syntactic signatures* along training trajectories reveals that they emerge during distillation and persist through subsequent preference optimization and reinforcement learning.
Chaeyeon Yi, Yun Am Seo
Tue 29 Sep
stat.APphysics.ao-ph
pdf html
- reading appear, critical slowing, consistent, slowing
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.72
- interest Macroeconometrics and time series 0.76
Abstract
Critical slowing down (CSD), the slowing of recovery from perturbations as a system approaches a critical transition, is widely used for early warning, but whether such signals appear in environmental transitions subject to rapidly changing external forcing has rarely been tested with observations. Here we show, using 2,427 icing-onset events in 1-min observations from 121 South Korean road weather stations, that such signals do not consistently precede road icing. Icing began from different prior surface states, but conditions at onset varied continuously with weak separation among pathways. Cold dose, deposition dose and freeze-thaw accumulation at onset exceeded those at non-icing control times, with cold-dose discrimination highest at decay timescales of 30 min or less. However, the increases in variance and lag-1 autocorrelation expected under CSD did not distinguish icing from control times; the same procedure recovered these increases in synthetic series with a declining recovery rate, and the null result was replicated for 3,091 events at 127 independent stations. Predictive information was nevertheless present: adding recent changes and cumulative antecedent conditions to current conditions improved discrimination, and surrounding meteorological fields increased the area under the ROC curve by 0.030-0.032 for forecast windows of 1 to 12 h and reduced false alarms by 62-67% at 3 and 6 h. These results suggest that road icing is more consistent with recent meteorological forcing moving the road surface into icing-favourable conditions than with a progressive decline in internal stability, and they indicate where early warning should look when a transition is externally forced.
Joris Pinkse
Thu 24 Sep
econ.EMstat.CO
pdf html
- reading jl, quadrature, rules, quadrature rules
- favorites near “Monte Carlo Geometry Processing” 0.78
- interest Computational methods for heterogeneous agent and dynamic economic models 0.73
Abstract
This paper contains new 80-digit positive-weight Gauss-Hermite and positive-weight interior-node Gauss-Legendre quadrature rules for up to five dimensions and varying polynomial degree accuracy (depending on quadrature type and dimension). Some of these rules improve on the best available rules in the literature and some offer rules where none (other than the tensor product) existed. The results were produced by combining methodology developed by previous researchers with new approaches. The full precision rules themselves are at https://doi.org/10.5281/zenodo.22881864. Software in Julia, Python, and R providing the rules is available via a package registry and/or GitHub: Quadriceps.jl (both 64-bit and 128-bit; https://github.com/NittanyLion/Quadriceps.jl), quadriceps-py (64-bit only; https://github.com/NittanyLion/quadriceps-py), and quadriceps-r (64-bit only; https://github.com/NittanyLion/quadriceps-r). The software used to create these rules is available via PositiveWeightQuadratureSolvers.jl (https://github.com/NittanyLion/PositiveWeightQuadratureSolvers.jl). The replication package is at QuadricepsReplicationPackage.jl (https://github.com/NittanyLion/QuadricepsReplicationPackage.jl). A snapshot of all five packages is archived at https://doi.org/10.5281/zenodo.22883240.
Marco Mandap
Mon 28 Sep
stat.MEcs.CLstat.AP
pdf html
- reading kalman, donsker, apparently, play
- interest Semiparametric inference and debiased machine learning 0.82
- your work near “Automated Solution of Heterogeneous Agent Models” 0.78
Abstract
We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count threshold. App-level rating histograms provide a distributional diagnostic for samples returned under different API sort orders; because star ratings are discrete, classical continuous Kolmogorov-Smirnov critical values are not used. A local-level state-space model and the Kalman filter provide a denoised temporal trend. Full proofs cover the BLUE and Gaussian maximum-likelihood result, Gaussian-conjugate shrinkage, the Glivenko-Cantelli and Donsker theorems, count transformations via the delta method, and exact Gaussian Kalman filtering. A worked three-review example shows how textual complaints can materially reduce an apparently perfect star-only score.
Adhithyan Kalaivanan, Zheng Zhao, Jens Sjölund, Fredrik Lindsten
Wed 30 Sep
cs.LGstat.ML
pdf html
- reading hilbert space, interacting, hilbert, smc
- favorites near “Monte Carlo Geometry Processing” 0.79
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.74
Abstract
Inference-time steering adapts pretrained diffusion and flow-based models to new tasks, e.g., to generate samples from a conditional distribution or samples with desired properties, without retraining. This can be formalized as sampling from a reward-tilted generative prior. As exact sampling from this distribution is intractable, guidance-based methods rely on approximations producing biased samples, and sequential Monte Carlo (SMC) methods correct for this bias using importance weights. However, while exact in the large particle limit, SMC suffers from weight degeneracy and particle collapse in practice. We propose interacting particle guidance (IPG), which replaces reweighting with transport. The particles interact through an additional drift, derived from the Feynman--Kac PDE to cancel the reweighting term, and remain unweighted. Choosing the drift in a reproducing kernel Hilbert space yields a closed-form solution that is cheap to compute, with negligible overhead compared to SMC. We demonstrate the method on Gaussian mixtures with known posteriors, and on high-dimensional image inpainting and protein structure inference tasks.
Joshua Chen, Peter Jan van Leeuwen
Wed 30 Sep
stat.MEstat.COstat.ML
pdf html
- interest Semiparametric inference and debiased machine learning 0.82
- favorites near “Variational Bayes under Model Misspecification” 0.85
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.77
Abstract
In Bayesian inference problems with non-Gaussian observation noise, the posterior is only as accurate as the noise density, and gradient-based samplers need that density and its gradient evaluable pointwise, whether from an explicit expression or from code, and without an inner solve. We propose Copula Active Subspaces (CAS) to represent this noise density. A componentwise rank transform isolates the noise law's dependence in its copula, and a rank-$r$ reduction keeps only the directions along which that dependence varies. These directions are the leading eigenvectors of the copula score covariance $\boldsymbol{C} := \mathrm{Cov}_{\pi_{\boldsymbol{Z}}}(\nabla\log c^{Z})$, which is what makes the reduction a copula active subspace. Because $\boldsymbol{C}$ vanishes when the coordinates are independent, these are directions of dependence, which the covariance of the data need not identify. From this construction follow a Gaussian-reference KL divergence bound with the explicit constant $\tfrac{1}{2}$, minimized over all rank-$r$ reductions by exactly this eigenspace; a diagnostic for the error the reduction leaves behind, computable from the samples alone; and, from Hermite score matching, a reduced log-density and gradient in closed form, with the truncation orders and the Stage-2 regularization constants chosen on validation samples. The reduction replaces a $d$-dimensional density estimation problem by an $r$-dimensional one. On a $d=20$ noise law and a Bayesian inference problem with that noise, CAS lowers noise KL divergence more than fivefold and posterior KL divergence more than sevenfold against Gaussian-copula, product-of-marginals, and PCA-subspace baselines, and lowers noise KL divergence by factors of about $3.5$ and $2.7$ on two further $d=20$ examples.
Jin Li, Ye Luo, Xiaowei Zhang
Fri 25 Sep
math.OCstat.MEstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.77
- reading theory, argument, future, mean-squared-error
Abstract
Adaptive algorithms increasingly make decisions while reshaping the dynamics that generate their future data. We establish a shrinking-tube concentration bound for projected stochastic approximation driven by an adaptive Markov chain. The bound guarantees, with high probability, that every iterate after a chosen time remains within a tolerance around the target that tightens over time. The probability of any exit after the chosen time admits a polynomially decaying upper bound, and a matching lower bound shows that its polynomial exponent cannot be improved in general under finite second moments. The result therefore identifies a sharp tradeoff between how quickly the tolerance shrinks and how rapidly the probability of any future exit decreases. We also extend the analysis to recursions with additional martingale-difference noise and predictable bias, showing how growth in the martingale-difference noise scale slows the decay of the exit-probability bound while predictable bias restricts the admissible tube shrinkage. The proof combines backward kernel replacement, a finite-time mean-squared-error bound, and a blockwise maximal first-exit argument. We apply the theory to inventory learning with stockout-dependent demand and fixed stockout costs, and quantify how numerical gradient accuracy affects the all-future reliability of the resulting policies.
Gatien Séguy, Thomas Rodet
Fri 25 Sep
cs.LGcs.AIstat.ML
pdf html
- reading operator, nominal, scale linearly, variational
- favorites near “Monte Carlo Geometry Processing” 0.79
- interest Kernels, operators and functional data 0.78
Abstract
Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior Sampling (DPS) and Pseudoinverse-Guided Diffusion Models ($\Pi$GDM), rely on scalar hyperparameters tuned per task, usually against the ground truth. We introduce FB-GDM, a fully-Bayesian guided diffusion method that removes this calibration step. Starting from the Gaussian approximation of $\Pi$GDM, we derive a closed-form conditional score that depends on two precision parameters (inverse variances), one associated with the denoising approximation and one with the observation likelihood, and treat them as latent variables inferred by variational inference at each reverse step. A separable factorization makes each update scale linearly with the number of pixels, so the inference stays tractable at full image resolution, at a cost comparable to one $\Pi$GDM run. FB-GDM requires neither the noise level nor the ground truth: its only inputs are the observation and the forward operator. Experiments on CelebA-HQ inverse problems establish two results. (i) The precision parameters, inferred from the observation alone, allow FB-GDM to outperform $\Pi$GDM at its nominal setting, even when the latter is given the true noise level, by up to 14 dB depending on the operator, and to match the ground-truth-calibrated $\Pi$GDM oracle within 0.1 dB. (ii) FB-GDM is robust when the forward operator, the noise level, or the image distribution changes: it stays close to a per-problem $\Pi$GDM oracle throughout and does not exhibit the hallucinations observed with DPS, whereas DPS substantially degrades at a fixed scale and $\Pi$GDM stays competitive only if it is re-tuned against the ground truth for each new problem. When the prior is applied to images outside its training set, this re-balancing between data and prior keeps FB-GDM faithful where a fixed face-prior guidance can otherwise hallucinate.
Haruka Ezoe, Ryohei Hisano
Mon 28 Sep
cs.CLstat.ML
pdf html
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.74
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.81
- reading operators, processes, makes, history
Abstract
Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or which usages support the attribution. We introduce Coupled Usage--Sense Processes (CUSP), which derives these answers from a single marginal preserving temporal process. A hierarchical coupling relates contextual distributions through latent usage components, while Markov composition makes adjacent and longer span correspondences compatible. Displacement operators quantify change magnitude and timing, split variation exactly between movement of component centers and reorganization within components, and attribute it to transported component pairs. Word-local modes resolve distinct directions of change and their activity over time, while representative passages from attributed components ground the analysis in text. Under a Gaussian mixture specialization, we prove parametric recovery of the operators and squared distances. Synthetic experiments support the predicted rate. CUSP remains competitive on English and German DWUG and recovers controlled Janus profiles while maintaining compositionally coherent transport. A large corpus of US court opinions demonstrates transition, mode, and passage attribution in unlabeled natural text. CUSP thus makes magnitude, timing, mechanism, movement, modes, and textual evidence compatible views of one lexical history.
Ganghua Wang, Shaddin Dughmi
Thu 24 Sep
stat.MLcs.LGmath.STstat.TH
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.80
- reading theory, question, striking, learning theory
- favorites near “Orthogonal Statistical Learning” 0.84
Abstract
This work revisits a fundamental question in active learning: how powerful is the ability to synthesize arbitrary queries? Compared to pool-based active learning, where the learner only selects queries from a given unlabeled pool, we find that this seemingly mild change in query ability may dramatically alter the difficulty of statistical learning. In particular, some hypothesis classes that are inherently slow to learn in the pool-based setting, achieving only polynomial error decay in the number of samples, become exponentially learnable once synthesized queries are allowed. This striking gap suggests that membership query synthesis induces a fundamentally different mode of learning, one that is not adequately captured by existing active learning theory and calls for new analytical tools to characterize its complexity. Motivated by this phenomenon, we develop several sufficient conditions, present intriguing examples, and propose a conjectural perspective toward understanding which hypothesis classes admit efficient learning through synthesized queries.
Carlos Baquero, Raquel Menezes
Tue 29 Sep
stat.AP
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.72
- interest Adaptive data collection and anytime-valid inference 0.80
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.72
Abstract
Power laws are often claimed for Bitcoin quantities from the slope of a log-log plot. We test one such claim with a formal protocol: whether the upper tail of the values of live unspent transaction outputs (UTXOs) follows a power law. We reconstruct the exact UTXO set at eight block heights from 2012 to 2025, validate it against an independent Bitcoin Core snapshot, and apply the Clauset-Shalizi-Newman protocol with three adaptations the data require: exact discrete support, several near-tied cutoffs retained and reselected in every bootstrap replicate, and resampling by transaction, because the outputs of one transaction are not independent. The protocol was fixed before inference and every later change is recorded. The pure power law is rejected at every height, with observed Kolmogorov-Smirnov distances five to forty times those the fitted law produces, and where the exponent is below 2 its extrapolation predicts, in expectation, several outputs each larger than the total coin supply. A lognormal, a stretched exponential and a power law with an exponential cutoff near a thousand bitcoin fit better in sample at every height, as families that contain the power law must, by margins that exceed those seen on simulated power-law tails at six of the eight heights, and they predict held-out transactions better at nearly every height. They rank differently from height to height, and the lognormal and the cutoff power law differ clearly at only two. The tail's shape and the share of supply it holds drifted slowly while the number of outputs grew seventy-fold. The rejection survives a null that reproduces the dependence measured within transactions, which halves the one exceptional effect size in the panel and changes no verdict. Bitcoin UTXO values are heavy-tailed but not power-law distributed.
Yves Tillé, María Guadarrama Sanz
Tue 29 Sep
stat.ME
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.82
- reading children, auxiliary, consistent, virtually
- interest Empirical Bayes, shrinkage and decision theory 0.76
Abstract
We propose a novel joint sampling design for surveys involving two interrelated populations of parents and children in the presence of complex family structures, including shared parental responsibility and blended families. The design combines a graph-theoretic representation of co-parenting relationships with balanced sampling. This framework ensures that, in the first sampling phase, exactly one parent is selected from each co-responsibility pair while maintaining accurate inference for both the parent and child populations. A second sampling phase selects one child per sampled parent using unequal probabilities that produce consistent weighting and unbiased estimation. The proposed framework accommodates overlapping family relationships while balancing auxiliary information for both populations. Results based on the Luxembourg population registers show that the estimated totals are virtually identical to the known population totals for all balancing variables, and other auxiliary variables, demonstrating the excellent precision and efficiency of the proposed sampling design.
Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki
Wed 30 Sep
cs.LGstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.79
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.79
- interest Adaptive data collection and anytime-valid inference 0.80
Abstract
Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length $N$; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed approximate conditional oracle. Under explicit regularity conditions and uniform Hellinger error bounds, for any fixed target accuracy $\varepsilon\in (0,1/8]$ and sufficiently large $N$, our sampler achieves seed-averaged total-variation error at most $\varepsilon$, with total masked-state submissions and sequential depth both bounded by $O(N^C \varepsilon^a)$ for constants $0 <C <1$ and $a > 0$. These guarantees use polynomial vocabulary size and an edge-response lower bound set by $N$ and $\varepsilon$. The sampler shares evaluations of hypothetical reveals across dependence tests to identify safe parallel batches without requiring full recovery of the hidden forest. A tunable parameter trades probing cost against irreversible commit rounds. In the same class, any admissible irreversible product-commit sampler attaining the same seed-averaged accuracy requires $\Omega(N^c \varepsilon^b)$ counterfactual submissions or commit rounds in the worst case, for constants $c,b>0$.
Yanjin Xiang, Zhihua Zhang
Thu 24 Sep
math.STmath.PRstat.MLstat.TH
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.78
- favorites near “Introduction to regularity structures” 0.81
- reading moment, matrices, random matrices, infinite
Abstract
We prove a strong-universality theorem at the exact fourth-moment threshold for finite families of non-Hermitian random matrices assembled from independent unordered-pair vector atoms. Within one atom, the matrix colors and the two endpoint orientations may have arbitrary joint real covariance, subject to reversal consistency across ordered type pairs and endpoint exchangeability in same-type blocks; the law may also depend on finitely many endpoint types. The matrices may be adjoined to an arbitrary deterministic tuple that converges jointly strongly with the type projections. For every fixed matrix amplification and fixed noncommutative \(*\)-polynomial, the resulting tuple converges strongly to an explicit covariance-matched free Gaussian family. Cross-type blocks are described by masked circular variables, whereas same-type blocks split into independent endpoint-symmetric and endpoint-antisymmetric semicircular sectors. Only a finite radial fourth moment is assumed off the diagonal, and a finite second moment suffices on the diagonal. The mode of convergence depends on the coupling: corners of one infinite array converge almost surely on a common event, while fixed-law nonnested triangular arrays converge in probability for each fixed test. As an application, we obtain exact-fourth-moment strong limits for finite-separable continuous left/right profiles, including separately weighted literal-transpose terms. For the nested almost-sure formulation, the fourth-moment threshold is sharp already on the Wigner subfamily; at this threshold arbitrary fresh rows admit no coupling-invariant almost-sure upgrade of our in-probability conclusion.
Beenu Sareena, Heike Hofmann
Mon 28 Sep
stat.CO
pdf html
- favorites near “Monte Carlo Geometry Processing” 0.83
- reading adjustment, highlight, functions, curves
- interest Kernels, operators and functional data 0.71
Abstract
Readable annotations are essential for interpreting statistical graphics, yet conventional single-color lines and labels can lose visibility when they cross backgrounds with heterogeneous luminance. We introduce a contrast-aware annotation framework implemented in the ggtwotone R package for ggplot2. The framework combines dual-stroke rendering, adaptive text-color selection, and perceptually guided highlight palettes to improve annotation visibility across light and dark regions. A shared contrast-adjustment engine supports WCAG- and APCA-based contrast criteria, reducing the need for manual color adjustment. The package provides contrast-aware geoms for segments, curves, paths, mathematical functions, regression overlays, and text while remaining compatible with standard ggplot2 workflows. A simulation-based evaluation across heterogeneous background colors demonstrates improved worst-case contrast for dual-stroke annotations and adaptive text selection, while also identifying trade-offs in perceptual separation as the number of requested highlight colors increases. Applications to statistical graphics and scientific images illustrate the framework in practical visualization settings. Together, these tools provide a reproducible approach for incorporating contrast considerations directly into graphical annotation.
Jinxiu Wen, Zhiming Li
Wed 30 Sep
stat.MEstat.AP
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.80
- interest Empirical Bayes, shrinkage and decision theory 0.81
- reading data analysis, empirical, control, frequentist
Abstract
Bilateral data from paired body parts are common and correlated in clinical studies. Classical frequentist methods perform poorly in small or sparse datasets. This paper develops a Bayesian framework for bilateral data with multiple groups under Dallal's model. We derive three objective priors (uniform, Jeffreys', and Bernardo's reference priors) and propose a range-based posterior testing procedure, combined with a decision rule, to test the homogeneity of risk differences, with the equivalence margin calibrated. Monte Carlo simulations evaluate empirical Type I error rates, powers, and interval estimation properties. Results show that the Bayesian methods achieve accurate coverage probabilities, narrower confidence intervals, and better Type I error control than the frequentist Wald approach, especially in small-sample and sparse-data settings. We illustrate the methodology with two real datasets. The proposed framework provides a robust and flexible tool for multi-group bilateral data analysis.
Zhang Yanhai
Tue 29 Sep
cs.LGcs.NEstat.ML
pdf html
- reading references, units, gains, written
- favorites near “Spectral State Space Models” 0.77
- your work near “Timing as an Action: Learning When to Observe and Act” 0.73
Abstract
Replay-based continual learning almost always consolidates in a dedicated offline phase or by interleaving replayed samples with the input stream, whereas brains also consolidate during wakefulness through local sleep, brief use-dependent off-periods of individual circuits. We ask whether a network trained by local, biologically constrained rules can consolidate with no offline phase at all. An isolation rule confines replay updates to hidden synapses invisible to the current input under k-winner-take-all dynamics, with optimiser state advanced only inside the mask; a refractory rotation rule makes units that have just fired sit out the next competition, widening the consolidable set; a homeostatic pressure and a relative-novelty gate decide when replay bursts fire and when rotation runs. This inverts the usual direction of non-interfering continual learning: the hidden computation on the current input is held invariant (exactly on the proven channels, and for all but 0.3% of waking samples per update elsewhere) while past memories are written into the degrees of freedom the current batch leaves unused. On class-incremental split-MNIST the system reaches 91.6+-0.3% with no offline phase, at or above the best offline-night schedule on two held-out splits, tied with DER++ and above experience replay, ER-ACE, A-GEM and unmasked local replay; in a single pass it leads DER++ (91.8% against 90.1%) while the night falls to 76.9%. The advantage is largest at small buffers and gives way to the backpropagation references at large ones; on split CIFAR-10 the system leads offline rehearsal and experience replay but trails ER-ACE and DER++. Rotation carries most of the gain; isolation adds the invariance guarantee. The mechanism is not tied to the local rule: under the same schedule a backpropagation network with k-WTA hidden layers gains from rotation, and isolation is again free on top of it.
Livia Popa, Sumanta Basu, Martin T. Wells
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.73
- interest Macroeconometrics and time series 0.78
- reading dimensionality, forecasting, science, dynamical
Abstract
Nonlinear dynamical systems with memory arise across science and engineering, yet uncertainty quantification for efficient forecasting methods such as Next Generation Reservoir Computing (NGRC) remains underdeveloped. We study Bayesian ridge and conformal prediction intervals for NGRC and characterize when their uncertainty estimates agree or differ. In low dimensions, their asymptotic widths are governed by different summaries of the residual distribution, so agreement depends on residual shape rather than dimensionality alone. In high dimensions, regularization introduces a further tradeoff between estimation variance, shrinkage bias, and posterior uncertainty, leading to an explicit transition between regimes where Bayesian intervals are wider or narrower than conformal intervals. We extend these results to quadratic NGRC feature maps and give sufficient conditions for transferring the analysis to temporally dependent forecast windows. Simulations and real-data experiments support the theoretical predictions and illustrate how residual distribution, regularization, dimensionality, and distribution shift affect interval calibration and efficiency. These results provide a principled framework for choosing and interpreting uncertainty quantification methods in reservoir-based forecasting.
Audra McMillan, Kunal Talwar, Felix Zhou
Wed 30 Sep
cs.LGcs.CRstat.ML
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.79
- reading viewed, gives, practice, learning algorithm
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.79
Abstract
Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.
Bingjing Tang, Shuangping Li, Julia A. Palacios
Fri 25 Sep
stat.MEq-bio.PE
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.78
- interest Bayesian computation and structural estimation 0.82
- reading posterior inference, state, view, inference
Abstract
The coalescent is a central framework in population genetics for modelling the ancestral relationships among sampled individuals through a genealogy, represented as a rooted and ranked binary tree. In this model, lineages coalesce at a rate inversely proportional to the effective population size, a time-varying quantity of primary interest. The bounded coalescent conditions genealogies on the time to the most recent common ancestor being bounded above by a fixed time. This model is useful in various contexts, such as phylodynamics of infectious diseases with known introduction times and single-cell lineage tracing in synthetic barcoding experiments. To our knowledge, there is no existing tool that infers variable effective population size trajectories under the bounded coalescent. We view estimation under the bounded coalescent as equivalent to estimation of the intensity function of an inhomogeneous point process. We provide an efficient algorithm for coalescent simulation under the bounded coalescent using point process methods, retaining the exactness of naive rejection sampling while substantially reducing computational cost and avoiding repeated numerical inversion of the bounded cumulative hazard. We then develop a Markov chain Monte Carlo procedure for posterior inference of effective population size trajectories that avoids discretization of the likelihood integrals. In simulations, conditioning on the bound reduces the median sum of squared errors in two of three settings, with less favourable results in the most rapidly varying setting. We illustrate the method using severe acute respiratory syndrome coronavirus 2 sequence data from Washington State.
Peter Vergeer
Tue 29 Sep
stat.AP
pdf html
- reading science, theory, evidence, main
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.79
- interest Language models as statistical objects 0.77
Abstract
In forensic science, when crime-scene evidence (CSE) and suspect-related evidence (SRE) is present, it is customary to report on the value of this evidence in the form of a likelihood ratio (LR). The LR can be calculated as the probability of CSE given SRE divided by the probability of CSE given that it was generated by a randomly selected person from an alternative culprit population. In forensic science, this is known as a feature-based LR, and intuitively the LR contrasts "similarity" by "typicality". Since it is generally a problem for feature-based LRs to find appropriate models for the data, one either resorts to score-based LRs or to adjusting the feature-based output post-hoc to well-calibrated output. Either way, the above definition of the LR is broken and interpretation of the LR as similarity between CSE and SRE divided by typicality of CSE is destroyed. Here, we report on progress in obtaining instantly well-performing feature-based LRs using gradient descent in combination with Bayesian probability theory to train a two-level model, the main LR model in forensic science for describing distributions of continuous data. For a dataset of laser-ablation inductively-coupled-plasma mass-spectrometry measurements on glass fragments from forensic casework, we show that our best model on validation data yields much better calibrated feature-based LRs on the test set when compared to state-of-the-art feature-based LR systems trained on the same type of data, and that it improves a factor of 4.5 on average on $C_{\mathrm{llr}}$. For LRs interpretable in terms of "similarity" contrasting "typicality", this is a major advancement. However, a state-of-the-art LR system still performs a factor of 1.5 better on $C_{\mathrm{llr}}$ for this data. We also present future plans to close this remaining gap. In order to facilitate collaboration, we have put relevant code on GitHub.
Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li
Fri 25 Sep
stat.MLcs.AIcs.LG
pdf html
- reading ex, expectations, language model, policy
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.78
- interest Language models as statistical objects 0.76
Abstract
Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a fixed-benchmark mean under an exact rollout budget. We develop Speculative Evaluation with a Hierarchical Bayesian Neyman (HBN) policy with pilot size and stage weight jointly chosen ex ante. It runs a short uniform pilot, pools per-task success counts with a hierarchical Bayesian model, and uses posterior expectations of task-level sampling variances for exact positive-integer Neyman allocation. To mitigate the pilot synchronization barrier, HBN-async speculatively executes continuations from partial pilot feedback and retains those selected by the final allocation. Across six checkpoints and 18 benchmark groups, we evaluate 107 nondegenerate benchmark-checkpoint profiles. For rollout budgets of 8-64 per task, Speculative Evaluation reduces variance relative to Uniform by 12.8%-33.6% on average across profiles, outperforming hindsight-tuned empirical and independent Bayesian baselines. Real-generation experiments that account for the pilot synchronization barrier show that HBN-async mitigates its overhead, helping translate statistical efficiency into practical evaluation benefits.
Ruichen Jiang, TaeHo Yoon
Fri 25 Sep
math.OCcs.LGstat.ML
pdf html
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.80
- interest Kernels, operators and functional data 0.79
- your work near “Computational Methods for Economic Models with Function Valued States” 0.78
Abstract
We study the deterministic oracle complexity of finding approximate solutions to composite monotone inclusion problems, formed by the sum of a smooth single-valued monotone operator and a maximally monotone set-valued operator, under the tangent-residual criterion. We introduce the Anchored Extra-Proximal (AEP) framework, which combines an anchored extrapolation step with an inexact anchored proximal update satisfying a relative-error condition. The framework recovers the composite Fast Extragradient method in the first-order setting and yields natural second- and higher-order extensions by replacing the operator in the implicit update with its Taylor approximation at the extrapolated point. For every $p\geq 2$, assuming that the $(p-1)$th derivative of the single-valued operator is Lipschitz continuous, we combine this construction with a bisection line search to obtain a $p$th-order method that finds a point with tangent residual at most $\varepsilon$ in $\widetilde{O}(\varepsilon^{-2/(3p-1)})$ oracle calls. This improves all prior upper bounds for $p$th-order methods: in particular, it improves the previous best-known $\widetilde{O}(\varepsilon^{-1/p})$ tangent-residual complexity as well as the classical $O(\varepsilon^{-2/(p+1)})$ bound of higher-order hybrid proximal extragradient methods under the weaker duality-gap criterion. We complement this result with a worst-case lower bound of $\Omega(\varepsilon^{-2/(3p-1)})$ for every deterministic algorithm in the $p$th-order oracle model, without restricting the algorithm to tensor steps or any other prescribed update structure. Thus, the proposed method attains the optimal dependence on $\varepsilon$, up to logarithmic factors, for all $p\geq2$.
Yutong Feng, Bowen Liao, See Kiong Ng, Yuxuan Liang
Tue 29 Sep
cs.LGcs.AIstat.ML
pdf html
- reading pretraining, forecasting, series, consistent
- favorites near “Spectral State Space Models” 0.76
- interest Macroeconometrics and time series 0.75
Abstract
Joint-embedding predictive architectures (JEPA) have emerged as a promising self-supervised pretraining paradigm for time series, learning representations by predicting target embeddings in latent space rather than reconstructing raw signals. Yet evidence on their benefits remains mixed, and most studies test only a single backbone or a narrow set of architectures, leaving unclear whether JEPA pretraining is a reliable improvement or one that depends heavily on the downstream model. We address this gap through a large scale evaluation of one JEPA instantiation across nine backbones and eleven benchmarks spanning temporal and spatio-temporal forecasting, the most extensive cross architecture assessment of JEPA for time series to date. We find that the benefit of this instantiation varies sharply across backbones, producing consistent gains for some architectures and consistent degradation for others, even on the same dataset. This pattern holds across both task families, indicating the variability is a general property of this instantiation rather than a dataset specific artifact worth accounting for when choosing a backbone in practice.
Damiano Brigo, Raphaël Huser, Dan Leonte
Mon 28 Sep
stat.MLcs.LGstat.APstat.COstat.OT
pdf html
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.72
- reading posterior inference, price, prices, functional
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.76
Abstract
Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can be propagated through heteroscedastic neural surrogate pricers for path-dependent exotic options. The resulting posterior-predictive distributions combine residual parameter uncertainty with conditional surrogate uncertainty and yield uncertainty-aware price intervals. We further introduce Hellinger-SHAP, an information-theoretic explainability method for posterior inference. Rather than attributing a single parameter point estimate, it applies local-background Kernel SHAP to a posterior-information functional measuring contraction from the prior to the posterior. This identifies maturity--moneyness regions associated with posterior information gain for individual rHeston parameters. In a simulation study, posterior-predictive intervals provide calibrated or conservative coverage across forward-start, barrier, and realized-variance claims, while point plug-in prices can be materially unreliable for selected contract regimes. Together, the UQ and XAI analyses provide a transparent framework for uncertainty-aware neural calibration and downstream exotic pricing under the specified prior-predictive model.
Seungjin Choi
Mon 28 Sep
stat.MLcs.LGstat.ME
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.79
- interest Language models as statistical objects 0.78
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.81
Abstract
Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled target inputs, allowing both the input distribution and its relationship with outcomes to change. We use Exponential Tilt Reweighting Alignment (ExTRA), introduced for classification by Maity et al. (2023), to estimate structured distribution shifts. We compare using its estimated weights in conformal calibration with additionally tilting the source predictive distribution. Shared learned predictors, estimated weights, calibration samples, and test observations isolate the effect of tilting. Existing theory gives both procedures target coverage with true weights and a common coverage bound with estimated weights. Identification calculations and an analysis of how scoring interacts with weight estimation error help explain why their performance can nevertheless differ. In a synthetic regression setting where the assumed models match the data-generating process and target inputs are informative about the shift, tilting reduces mean set length by about $30\%$ relative to weighting alone, with both methods attaining coverage near nominal. Tilting can instead cause substantial coverage losses in synthetic classification and in regression when target inputs provide little information about the response shift. Real-data experiments also show no consistent benefit. Good coverage from weighted calibration alone does not ensure that adding predictive tilting will preserve coverage. Deciding when to apply this additional adjustment using only source labels and target inputs remains an open problem.
Shivam, Bhargab Chattopadhyay, Nil Kamal Hazra
Thu 24 Sep
stat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.81
- reading econometrics, econometric, procedures, burden
- interest Adaptive data collection and anytime-valid inference 0.80
- looks like clinical and biomedical studies
Abstract
We study the problem of selecting the extreme (best or worst) population from among $K(\geq 2)$ populations, under the assumption that the extreme population is sufficiently separated from the nearest population. The selection is based on an appropriate measure, which may vary across different application domains. Since the actual value of the measure is unknown, we obtain its estimator using the generalized method of moments under a multistage sampling design. Using this estimator, we propose two sequential procedures, namely online algorithm and multi armed bandit based algorithm. Under suitable regularity conditions and without imposing parametric assumptions on the underlying distributions, both algorithms correctly identify the extreme population with a desired level of confidence. We illustrate the proposed algorithms through applications in econometrics and genetics. In the econometric application, the extreme population is selected using the Gini index as a measure of inequality and the performance of the proposed procedures is assessed through extensive Monte Carlo simulation studies conducted under various distributional settings. In the genetics application, the worst population is identified using a measure derived from the tumor mutation burden (TMB) score and the practical applicability of the proposed algorithms is demonstrated using the Memorial Sloan Kettering-IMPACT 50000 clinical sequencing cohort. Further, we use the proposed framework to identify an anomalous population, provided such a population exists.
Sangsidhya Kar
Wed 30 Sep
cs.LGstat.ML
pdf html
- reading equilibrium, language model, generating, tends
- interest Language models as statistical objects 0.79
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.77
Abstract
In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps when it is the target. SGD simply adds these contributions. Coordinate-wise adaptive methods such as Adam, RMSProp, and sign descent instead divide each update by a running estimate of its magnitude, and that estimate is largest immediately after the token appears. This imbalance has two effects. At the level of the whole output layer, we characterize which optimizers preserve the mean output embedding: every method whose update is linear in past gradients does, as do Kronecker-factored and orthogonalized methods such as Shampoo and Muon. Adam, Adafactor, Lion, and sign descent do not, and for these methods we obtain an exact step-by-step expression for the change. At the level of an individual rare token, the same normalization shifts the training fixed point. In the unigram model, sign descent lowers the logit of every token that occurs in fewer than half of the minibatches at a constant expected rate. For RMSProp with periodic arrivals, we can solve the fixed point in closed form: if a token is absent for at least two consecutive minibatches, its equilibrium probability is strictly below its data frequency for every learning rate, and the ratio tends to $\kappa/(2(e^{\kappa/2}-1))$. Here $\kappa$ is the mean number of steps between occurrences divided by the second-moment time constant $1/(1-\beta_2)$. In the same model, SGD and AMSGrad retain the unbiased fixed point. We test these predictions both in a unigram model and in a small language model trained from a known generating distribution. With random arrivals, the bias is larger than the periodic formula predicts; in the language model, the optimizers with the biased fixed point also fit the generating distribution less well.
Stina Zetterstrom, David S. Robertson, Sofía S. Villar ( MRC Biostatistics Unit, University of Cambridge, United Kingdom)
Wed 30 Sep
stat.COstat.AP
pdf html
- reading software, practice, reviewed, available
- your work near “Timing as an Action: Learning When to Observe and Act” 0.81
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.84
- looks like clinical and biomedical studies
Abstract
Response-adaptive randomization (RAR) modifies treatment allocation probabilities during a clinical trial as response/outcome data accumulate, with the aim of improving patient benefit, statistical efficiency, or both. Despite substantial methodological development, adoption of RAR in clinical practice has remained limited, and the software available to support its design and implementation has not previously been reviewed. We identified 16 publicly available, open-source software packages implementing RAR methods, spanning urn-based, target-allocation, Bayesian, Markov decision process (MDP)-based, and dose-finding approaches, and evaluated them with respect to their methodological and practical characteristics. We found that while several software packages exist, most are method-specific, and only a small number provide broader, general-purpose adaptive-trial design and analysis capabilities. Explicit support for practically relevant features, including delayed and missing outcome data, temporal trends in response rates, flexible operating-characteristic evaluation, and platform or multi-arm multi-stage trial designs, is rare or absent across the identified software. These findings suggest that while a diverse set of tools exists for exploring RAR designs, gaps remain between the methodological literature and the software available to implement it in practice. Continued development of flexible, practically oriented, and validated software is important for the wider adoption of RAR in clinical research, and we highlight interesting areas of further work.
Mehrdad Pournaderi
Tue 29 Sep
stat.MLcs.LGmath.STstat.MEstat.TH
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.77
- favorites near “Orthogonal Statistical Learning” 0.85
- interest Online learning and statistical learning theory 0.81
Abstract
Weighted split conformal prediction reweights calibration scores by the likelihood ratio between the test and training covariate distributions and guarantees marginal coverage under covariate shift. We study its coverage conditional on the calibration data. An elementary argument, based on a single concentration inequality at a fixed population quantile, gives explicit training-conditional bounds without unspecified constants, and shows that the relevant scale is not the supremum of the likelihood ratio but a variance proxy built from the chi-squared divergence of the shift and from the average of the ratio over the part of the test population, of probability equal to the miscoverage level, where it is largest. A two-point lower bound shows that the root-m rate and the chi-squared contribution are intrinsic to the shift. Run at an explicitly inflated level, the weighted quantile becomes a deterministic PAC prediction set. We compare it with randomized rejection sampling and with importance-weighted learn-then-test and, through a certified choice of a clipping level for the likelihood ratio, map the regime in which each gives the narrower valid set. The analysis extends to estimated likelihood ratios and to tail functionals estimated from an unlabeled source sample, which yields a fully finite-sample certificate.
Yuanhao Chen, Peter Chin
Fri 25 Sep
cs.CLcs.LGcs.SDstat.AP
pdf html
- reading audio, language, languages, pairs
- interest Language models as statistical objects 0.80
- favorites near “Spectral State Space Models” 0.74
Abstract
Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members' mean representations. Averaging those offsets gives a direction for each stream, and we measure the cosine between the two. Because the two streams already agree about arbitrary phoneme pairs, we compare every measure against a reference built from random pairings rather than against zero. We apply this to 6 models, 7 features and 15 languages from 11 families. Only voicing in the two Qwen2.5-Omni models exceeds that reference after correction for multiple testing, and the reference varies by a factor of seven between models. In three of the six models, voicing has one direction in audio across the 14 languages with enough minimal pairs to measure it, and every language pair agrees in two of them. The model family, not the model size, predicts which stream represents a feature.
Md Musfiqur Rahman, Kenneth Lee, Ziwei Jiang, Padmaja Jonnalagedda, Ruocheng Guo, Murat Kocaoglu
Wed 30 Sep
cs.LGcs.AIstat.ML
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.78
- reading causal, root, complicated, online
- interest Language models as statistical objects 0.76
Abstract
Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, existing approaches rely either on heuristic methods or on conditional independence tests with a strong unconfoundedness assumption, and thus fail to exploit other complicated distributional constraints in the presence of latent variables. To relax these assumptions, we model the underlying system as a causal model and the anomalous system as a change in the structural functions of the same causal model. Specifically, to handle unobserved confounders, we establish an implicit connection between distributional constraint testing and root cause analysis. To adapt our approach to data generated from arbitrary causal models, we employ the deep causal model (DCM) framework, in which we design the causal model using neural networks. Finally, we illustrate how our method, RCA-DCM, can utilize different levels of partial graphical knowledge to perform RCA. We evaluate RCA-DCM against state-of-the-art baselines on simulated datasets, a physics-based causal chamber and two micro-service applications. RCA-DCM improves top-1 accuracy over the strongest baseline on both Sock Shop (0.880 vs. 0.752) and Online Boutique (0.776 vs. 0.712), and when the true root cause in the causal chamber is unobserved and acts as a latent confounder, it recovers the exact root-cause set more often than any competing method (perfect recovery rate (PRR) 0.846 vs. 0.731).
Kirk Bansak, Jens Hainmueller, Dominik Hangartner, Jeremy Ferwerda, Elisabeth Paulson, Angie Delevoye, et al.
Tue 29 Sep
econ.GNcs.CYq-fin.ECstat.AP
pdf html
- reading employment, labor, gains, markets
- favorites near “Why Do People Stay Poor?” 0.77
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.69
Abstract
Refugee integration is a central policy challenge for host countries, and where governments initially place refugees shapes their integration trajectories. Yet placement officers often have limited information about where each case is most likely to succeed. Algorithmic refugee matching uses administrative data, machine learning, and constrained optimization to recommend employment-optimized placements in real time as cases arrive, with human placement officers retaining final authority. Between January 2020 and June 2023, the Swiss State Secretariat for Migration randomly assigned about 2,000 refugee cases to receive a canton recommendation either algorithmically optimized for employment or drawn to approximate existing procedures, with placement officers and refugees blinded to assignment. The two arms used identical but separate canton and origin-group quotas, so gains reflect better refugee-canton matching rather than reallocation toward stronger labor markets. The trial began just before the COVID-19 pandemic shifted labor-market conditions. For the pre-registered primary outcome -- the share of months employed during the first three years -- the pooled intention-to-treat (ITT) estimate across the 2020-2023 placement cohorts was +2.2 percentage points (about 10% of the 22.3% control mean; 95% CI [+0.05, +4.33]), rising to +3.9 pp (about 17%; [+1.11, +6.68]) for the post-COVID 2022-2023 cohorts. Effects grew over time: at 36 months, the pooled ITT on the employment rate was +5.2 pp (about 11%; 95% CI [+1.10, +9.25]) -- comparable to the gains from hundreds of hours of intensive language training. Overall, the results provide rare field evidence that AI-based decision support can improve high-stakes public-sector allocation, offering a scalable, low-cost way to raise refugee employment.
Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski
Mon 28 Sep
stat.MEcs.AIcs.LGmath.STstat.TH
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.80
- your work near “Valid Inference with Imperfect Synthetic Data” 0.76
- reading goal, interventions, possible, generation tasks
Abstract
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.
Thomas B. Michelon, Hsieh Fushing
Mon 28 Sep
stat.CO
pdf html
- reading data analysis, good, science, interacting
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.75
- favorites near “Data Science at the Singularity” 0.81
Abstract
In this A.I. era, Computational Taxonomy is proposed to study complex systems by analyzing their databases under taxonomic hierarchies abiding the Principle of Science by providing "good explanations". Comparisons among branches or classes are carried out by Scientific Data Analysis (SDA) paradigm that explores all potential associative patterns, including interacting effects of all high orders, and evaluate finite sample precisions for all information pieces individually by effectively making use of all variables' categorical nature. Under each comparison, all confirmed information pieces are collected and displayed along row-axis of a heatmap with all involved study-subjects on the column-axis. Each comparison's heatmap individually characterizes participating classes and study-subjects and simultaneously provides a scientific basis for outlier detection upon all non-participants. All these heatmaps then collectively constitutes so-called Class-informatics that offers good explanations based on characteristic of all classes and study-subjects. Computational Taxonomy's Class-informatics indeed resolves multiple fundamental issues: Tukey's more than 60 years outlier detection problem, issue of self-correction annotation, and a crucial check on assumption of information-content equality between testing and training data sets in Machine Learning. A showcase of Computational Taxonomy is exclusively illustrated on Iris data.
Katharine Fisher, Michael Herbst, James Kermode, Youssef Marzouk
Mon 28 Sep
physics.chem-phstat.ML
pdf html
- favorites near “Monte Carlo Geometry Processing” 0.81
- reading science, computations, conformal methods, multistage
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.73
Abstract
Machine learning has become the standard tool for the design of interatomic potentials which balance efficiency and accuracy, but uncertainty quantification remains an open problem. Multiscale simulations introduce an additional challenge: robust uncertainty quantification across scales. Even within one scale, computations are often multistage, producing a sequence of target quantities, each dependent on the previous, and each with some uncertainty. Conformal methods have emerged as a model agnostic framework for recalibrating surrogate predictions to produce sets which contain the truth at a user-specified rate. For multistage workflows, we require uncertainty calibration for multiple chemical properties and atomistic configurations, and we want to propagate uncertainty sets to downstream quantities of interest. Such propagation should capture the error cancellations which occur in many downstream targets in materials science; for instance, an approximate energy difference is often more accurate than individual energy predictions. We present the first exploration of multivariate conformal methods for chemical properties, including Bonferroni-corrected hyperrectangles, hyperellipsoidal sets based on the Mahalanobis distance, and custom loss functions within conformal risk control. Calibration is applied directly to predicted energies, atomic forces, and virial stresses, then propagated to elastic constants and vacancy formation energies employing a variety of commonly considered approximate protocols in materials modeling. We highlight the benefits of building correlation predictions into the conformal procedure, making it possible to build sets which capture near symmetries and error cancellation. We conclude with a discussion of the interplay of the employed approximate computational protocol and conformal guarantees.
Jing Jia, Sifan Liu, Guanyang Wang
Tue 29 Sep
cs.LGcs.CVstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- reading unconditional, keeping, cheaply, closely
- favorites near “Spectral State Space Models” 0.75
Abstract
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.
Yuwen Qian, Yidong Ouyang, Zhengyan Wan, Hongyuan Zha
Wed 30 Sep
cs.LGstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.80
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.77
- interest Bayesian computation and structural estimation 0.79
Abstract
Sampling from unnormalized distributions over large discrete state spaces becomes difficult when a multimodal target is far from a tractable reference. We introduce Iterative Exact Discrete Guidance (IEDG), a population-exact, trajectory-wise guidance framework for unnormalized discrete targets. Rather than learn the full reference-to-target correction in one step, IEDG introduces a global Boltzmann tilt along an annealing trajectory. Each stage learns a stage-local posterior correction for an incremental Boltzmann tilt of the current source, while the resulting corrections are accumulated relative to a fixed analytic posterior. At the population optimum, exact stage posteriors recover the correct reverse dynamics, whose exact simulation reproduces the target distribution. IEDG chooses stage increments by relative effective sample size (rESS), which controls R\'enyi-2 displacement and locally adapts the step size to the thermodynamic geometry of the annealing path. Our stagewise total-variation analysis shows that limited overlap amplifies Bregman fitting error by $1/\sqrt{\mathrm{rESS}}$, while posterior, simulation, and truncation errors enter additively. IEDG improves all distribution-level errors over the neural baselines on ordered, exactly enumerated Ising $4\times4$, while substantially reducing one-shot errors on Ising/Potts $16\times16$ across thermodynamic regimes and attaining the best neural-sampler result on several reported local-statistic and phase-coverage metrics. On Max-Cut, its best-of-512 and average-sample ratios exceed all the baselines. Code and artifacts are available at https://github.com/StillFantasy123/iterative-exact-discrete-guidance.
Abhijit Brahme, Ishan Mehta, Gregory J. Matthews, Alexander Franks
Tue 29 Sep
stat.AP
pdf html
- reading continuous time, historically, american, advances
- interest Bayesian computation and structural estimation 0.78
- favorites near “Recent Developments in Machine Learning Methods for Stochastic Control and Games” 0.74
Abstract
Historically, statistical analysis of offensive lineman has been hindered by the lack of easily measurable quantities. More recently, with the introduction of player tracking data new methodological advances are now possible. Using high-dimensional spatio-temporal data, we adapt the defensive-matchup hidden Markov model of \cite{franks2015characterizing} from basketball to football pass protection, producing frame-by-frame probabilistic assignments of each pass blocker to the rushers. We show how this probabilistic assignment is a usable modeling artifact that augments existing player-evaluation frameworks. We directly quantify the attention a rusher commands, upgrade adjusted plus-minus \citep{Macdonald+2012} from all-or-nothing stints to partial, continuous blocking credit in continuous time, yield block-shedding survival metrics, and measure the space a rusher generates for his teammates. Fit to the first eight weeks of the 2021 NFL season, the resulting metrics recover widely-recognized elite rushers and pass protectors and align with independent charting.
Vincent Uhse
Wed 30 Sep
cs.LGstat.ML
pdf html
- reading retrieval, gives, forecasting, past
- favorites near “Learning Scalable Deep Kernels with Recurrent Structure” 0.78
- interest Language models as statistical objects 0.75
Abstract
Deep forecasters predict from a fixed-length lookback window, and lengthening it gives diminishing returns at a growing cost. Retrieval augmentation instead shows the model how similar past situations continued. Retrieving a whole past window gives every variate the continuation of the same past moment. In multivariate series, however, the best past match differs from variate to variate. We present GNA (Granular Neighbor Assembly), a retrieval layer for forecasting backbones that assembles neighbors at two granularities: whole past windows, which keep the variates coherent, and per-variate neighbors, in which each variate takes its future from its own best-matching past. A learned gate decides, per forecast step and variate, how much to trust these futures against a persistence forecast, next to the backbone's own forecast. Candidates come from an embedding trained to predict each window's future, and retrieval is strictly causal: a past window is used only once its future has been observed. With the same lookback for every model and the same retrieval constants for all datasets, GNA improves two Transformer backbones in 85 of 96 dataset-horizon settings, gives the lowest MSE on 8 of 12 standard benchmarks and beats its backbone in every seed on 10 of them. Both granularities are needed, and neighbors of mismatched queries are worse than none. Retrieval helps most where the lookback says least: the gate shifts trust to retrieved futures further ahead. Where it fails, on hourly non-stationary series at long horizons, the loss is consistent with a drifting level of the retrieved futures.
Jintao Wei, Chenxi Li, Songhao Wang
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.82
- interest Reinforcement learning, policy learning and off-policy evaluation 0.76
- favorites near “Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes” 0.79
Abstract
Federated Bayesian Optimization (FBO) enables distributed agents to collaboratively optimize expensive black-box objectives without sharing raw local observations. However, effective knowledge transfer remains challenging under communication constraints and task heterogeneity. We propose GUIDE-FBO, in which agents exchange compact distributions over the locations of their respective optima inferred from local Gaussian process (GP) posteriors, rather than raw observations, query points, or surrogate parameters. The server merges and reweights these distributional components before returning a subset to each agent. Each agent then constructs a Federated Interventional GP (FI-GP), which preserves the local posterior mean and spatially rescales its covariance for local decision making. For the upper confidence bound (UCB) instantiation, GUIDE-UCB, we prove that any bounded FI-GP uncertainty intervention preserves the leading-order cumulative regret rate of standard GP-UCB. When the transferred distributions place greater support near an optimum than in a suboptimal region, selecting the latter requires greater local posterior uncertainty. Experiments on 12 synthetic benchmarks and three real-world optimization tasks show that GUIDE-FBO remains effective across settings ranging from homogeneous to severely heterogeneous. Ablation results highlight the importance of spatially localized uncertainty intervention, while the communication analysis shows that GUIDE-FBO exchanges only compact distributional messages.
Martin Bladt
Wed 30 Sep
stat.ME
pdf html
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.78
- interest Causal inference and identification in econometrics 0.77
- favorites near “Optimal uniform convergence rates and asymptotic normality for series estimators under weak dependence and weak conditions” 0.81
Abstract
We study tail dependence of a $d$-dimensional random vector whose coordinates are subject to random right censoring. Along each fixed direction the censored problem reduces exactly to a univariate one, and the observed data determine the radius and whether it was produced by the event or censoring vector. An ordinary Kaplan--Meier product limit therefore estimates the joint tail probability in that direction, in every dimension, and with no multivariate survival function, no smoothing and no tuning parameter beyond the threshold. The theory of this directional estimator is formulated under an arbitrary marginal standardization and conditions only imposed on the standardized laws, in particular for any max-domain of attraction. We prove uniform consistency and functional weak convergence at the square root of the effective number of joint extremes, allowing the standardization to be estimated. A multiplicative standardization recovers the heavy-tailed theory, whereas a standardization built from the marginal (non-directional) Kaplan--Meier estimators requires no marginal tail model and targets the normalized tail copula itself. The standardization error is negligible for the former, and for the latter under a mild condition on the joint censoring. Simulation studies validate the finite-sample performance of the estimator. An application to the National Flood Insurance Program dataset comprised of claims generated by Hurricane Ian estimates the joint upper tail of building and contents losses from indemnities, which are subject to capping, simultaneously in both coordinates.
Xiaoran Cheng, Sen Na, Jia Li
Tue 29 Sep
cs.LGcs.AIstat.MEstat.ML
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.82
- interest Language models as statistical objects 0.81
- your work near “Automated Solution of Heterogeneous Agent Models” 0.74
Abstract
Deep ReLU networks are piecewise-affine mappings that partition the input space into cells, each characterized by a distinct activation pattern. This structure motivates fitting a local linear model within each cell to preserve predictive accuracy while improving interpretability. The challenge is to identify regimes that are stable, data-adaptive, and easy to explain. We propose REALM, a mixture of linear models whose regimes are induced by neural activation patterns. Because the number of activation cells in a deep neural network (DNN) can grow rapidly with depth, we first distill a deep teacher into a wide, shallow student network (WSSN), then binarize and cluster its hidden-layer activations to define the regimes and fit a linear model within each regime. Since the regimes are discovered from internal structure, the router does not carry the predictive burden. To make regime assignment interpretable, we train a multiclass logistic regression, the explanatory gate, to reproduce the regime assignments. The two-level structure is interpretable at both stages in terms of raw tabular or learned convolutional features: the gate identifies features that determine regime assignments, while the linear models identify features that drive predictions within each regime. We analyze an idealized setting that illustrates a trade-off between partition complexity and stability: as the number of regimes grows, finer partitions can improve approximation but may reduce regime-assignment stability. Experiments on tabular and image datasets show that REALM achieves competitive predictive performance relative to other DNN-guided mixture surrogates and inherently interpretable models while producing stable regime-level explanations.
Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta
Mon 28 Sep
cond-mat.dis-nncs.LGstat.ML
pdf html
- reading empirical, generalization, spectrum, committee
- interest Kernels, operators and functional data 0.78
- favorites near “A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning” 0.84
Abstract
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and the Hessian spectrum are analytically tractable in the proportional high-dimensional limit. First, we use a zero-temperature Gibbs formulation to obtain predictions for the observables of the typical minimizers of the empirical loss. Secondly, we use Edwards-Jones formalism to derive the limiting Hessian resolvent around these typical minimizers. All predictions agree with finite-size gradient-descent simulations. Finally, we study three measures of flatness, namely the left and right edges and the spectral mean, and check if a decrease in generalization error as the dataset size is increased corresponds to an increase in flatness. We find that the answer strongly depends on the learning task and on the ratio of the number of parameters to the number of data points. In regression, the spectral mean and right edge correlate with the generalization error, while the left edge does so only in the overparametrized regime. In classification this correlation reliably holds only in the highly overparametrized phase, while for underparametrized networks it can even reverse.
Yi Ding, Lan Wei, Xuehu Zhu, Wenlin Dai
Tue 29 Sep
math.STstat.TH
pdf html
- reading sequential, martingales, gains, online
- interest Adaptive data collection and anytime-valid inference 0.80
- favorites near “Eigendecompositions of Transfer Operators in Reproducing Kernel Hilbert Spaces” 0.80
Abstract
E-processes provide a flexible framework for anytime-valid inference, with the conventional rejection boundary $1/\alpha$ typically justified by Ville's inequality. Such a boundary is universal but can be conservative, as it does not exploit additional information about the null distribution. We propose a reference-null calibration framework that uses independent null samples to construct sharper rejection thresholds while preserving type-I error control. To study the statistical gain from sharper thresholds, we specialize the framework to conformal martingales. Building on histogram-based betting, we incorporate Krichevsky--Trofimov smoothing and develop a restart-mixture construction for distribution shift detection. We establish quantitative results for detection power and detection delay and characterize how the reduction in the rejection boundary translates into power and delay gains relative to the conventional Ville boundary. Numerical experiments corroborate the theoretical findings across a range of distributional changes. Finally, we apply the proposed calibration strategy to an existing e-process for online LLM watermark detection, showing that the method can improve sequential detection performance without modifying the underlying e-process. These results demonstrate that reference-null calibration provides a general and modular way to enhance the efficiency of e-process-based sequential inference.
Binh Thuan Tran, Khai Nguyen
Tue 29 Sep
stat.MLcs.LGmath.STstat.COstat.TH
pdf html
- favorites near “Monte Carlo Geometry Processing” 0.79
- interest Kernels, operators and functional data 0.79
- your work near “Computational Methods for Economic Models with Function Valued States” 0.77
Abstract
We propose sliced Orlicz-Wasserstein (SOW) distance which is a generalization of sliced Wasserstein (SW) distance. SOW replaces the $L^p$ norm in SW with a Luxemburg norm cost induced by an Orlicz function $\phi$. First, we prove that SOW distance is a metric on the space of measures with finite Orlicz norm, and show that it recovers the SW distance when the Orlicz function is $\phi(x)=x^p$. Next, we derive the topological properties of the SOW distance. In particular, we show that convergence under SOW implies weak convergence, and the converse is true under the compact support condition. We then present the theoretical results for estimating the SOW distance. We derive sample complexity for both the distance itself and the powered functional of the distance, and prove their minimax optimality. In addition, we discuss the computational algorithm for approximating the SOW distance by Monte-Carlo estimation and bisection search, as well as the associated approximation error and computational complexity analysis. Our experimental results reveal the superior computational efficiency of SOW compared with Orlicz-Wasserstein (OW) distance. Also, in the experiments, we demonstrate the favorable flexibility of SOW distance over SW in detecting differences between distributions by comparing their performance in two-sample tests and evaluating generative models on image datasets.
Jinhyeok Park, Hyeyoung Maeng, Hoseung Song
Mon 28 Sep
stat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- interest Macroeconometrics and time series 0.79
- reading structural, nonparametric, procedures, favorable
Abstract
We study nonparametric multiple change-point detection for high-dimensional sequences, aiming to identify time points at which the underlying distribution changes. While many existing methods perform well when change-points are well-separated, their performance can deteriorate when structural breaks are densely clustered. To address this challenge, we propose gBottomup, a graph-based bottom-up framework for multiple change-point detection in high-dimensional settings. gBottomup constructs a hierarchical segmentation by proposing merges of adjacent segments and verifying them through an unmerge rule that combines absolute significance with relative local heterogeneity, thereby adaptively refining partitions and estimating the number of change-points. Simulation results demonstrate that gBottomup performs reliably across a range of structural configurations and is particularly effective in frequent change-point settings, where existing top-down procedures may lose sensitivity. Runtime experiments indicate favorable computational performance relative to graph-based top-down alternatives. We illustrate the proposed method through an analysis of a S&P 500 dataset.
Bradley Huynh, Moyi Tian, Nancy Rodríguez
Wed 30 Sep
cs.SIcs.CYphysics.soc-phstat.AP
pdf html
- reading online, empirical, aggregate, spreading
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.73
- interest Language models as statistical objects 0.74
Abstract
Social media increasingly shapes everyday life, serving as both a central venue for discussion of major events and a space where online collective behavior can spill over into real-world activity, while artificial intelligence (AI) is likewise becoming increasingly influential across society. Understanding how different online communities respond to AI-related events, and disentangling the mechanisms underlying the development of such engagement, are therefore increasingly important for studying information spreading and the societal reception of AI. Our work addresses the limited connection between empirical studies of AI-related online engagement and mathematical modeling of group-level spreading dynamics. We develop a framework to collect and organize empirical Bluesky activity into distinct user groups, then use a network-based dynamics model to investigate the mechanisms underlying their engagement. We use the release of DeepSeek R1 as a case study under two complementary grouping schemes: AI-related communities and academic disciplines. For each group, we fit a network-based Susceptible-Infected-Recovered (SIR)-type model augmented with an exogenous engagement term, allowing us to quantify the relative strengths of endogenous network-driven spreading and direct external response. Across groups, we find that strong direct responses to the event do not necessarily coincide with strong network-driven propagation, revealing distinct engagement patterns that may be obscured by aggregate activity alone.
Rares Dimitrie Grozavescu, Mark Girolami
Wed 30 Sep
stat.MEmath.STstat.MLstat.TH
pdf html
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.79
- interest Semiparametric inference and debiased machine learning 0.81
- favorites near “Optimal Decision Rules for Weak GMM” 0.83
Abstract
Basis-restricted priors in Bayesian inverse problems can lose coverage when the truth has components outside the basis. We show that estimating the observation-noise variance can hide this loss. Under a linear forward model, when the in-span prior variance dominates the noise, the maximum-likelihood noise estimate absorbs the out-of-basis energy in the complement of the model range. Residual-magnitude and observation-coverage checks then stay near nominal while field coverage falls. We study repeated problems sharing one forward operator and one basis, fixed independently of the tested data. After projection onto the complement, and conditionally on the fitted noise scale, every exact test is a test of the scale-free direction of the residuals. We test the shape of their sample spectrum with John's sphericity statistic. Under Gaussian noise its null model is exact at finite sample size, and we derive its null mean and its power at proportional dimension. On synthetic problems and in a preregistered GEBCO topography study, the test detects structured out-of-basis variation that cross-validation and observation-coverage checks largely miss. It cannot detect Gaussian out-of-basis variation that is isotropic in the complement, since that is indistinguishable from a change of noise scale.
Sungwoo Park, Sunghee Park, Won Chang
Tue 29 Sep
stat.MLcs.LG
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.76
- reading foundation models, heterogeneity, inference, predictive distributions
- interest Semiparametric inference and debiased machine learning 0.80
Abstract
Tabular foundation models (TFMs) provide predictive distributions for regression, but their prediction regions can exhibit undercoverage or overcoverage even when point predictions are accurate. We introduce C-USIM (Conditionally-Uniformized Score Integration Method), a lightweight application of highest predictive density split conformal prediction that accommodates multimodal predictions. Given calibration and test outputs, it requires no additional training or model inference. It provides finite-sample marginal validity under our assumptions. We bound conditional-marginal coverage gaps using distribution-estimation error and score discreteness, and examine coverage heterogeneity through percentile rank-score plots. Experiments with TabPFN and TabICL show improved marginal coverage accuracy and lower average conditional and group coverage errors. Under a fixed data budget, allocating more observations to calibration can reduce marginal coverage error despite less accurate point predictions.
Md Khalid Hasan Sakib, Dristi Datta, Manoranjan Paul, Davina White
Fri 25 Sep
stat.APcs.LG
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.83
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.66
- interest Empirical Bayes, shrinkage and decision theory 0.75
Abstract
Reliable intervention-oriented uncertainty quantification from Earth observation (EO) remains challenging when continuous treatment shifts, spatial dependence, limited support, and satellite-outcome uncertainty must be addressed simultaneously. Existing causal, conformal, and spatial approaches address parts of this problem, but their direct combination does not generally recover the appropriate interventional reference law because candidate reassignment jointly alters treatment likelihood, standardized residuals, and graph-dependent residual likelihood. This study presents GeoDose-CP, a support-aware conformal framework for localized stochastic potential outcomes under continuous or mixed continuous-atomic treatment. Its central methodological contribution is a graph-local target-orbit law that jointly represents intervention-induced treatment shift, the inverse outcome-scale Jacobian, and spatial residual dependence. The framework further provides exact weighted candidate inversion, a scalable sparse approximation with explicit discrepancy accounting, and refusal under inadequate support. Evaluation used controlled known-truth experiments, MineDoseBench, treatment-density sensitivity analysis, external conformal comparators, and a multi-mine New South Wales (NSW) study. In MineDoseBench, GeoDose-CP achieved mean selective coverage of 0.9692 across 27 configurations and a minimum local q0.05 of 0.8951; exact-sparse auditing produced nine inclusion disagreements over 2,700 targets. In the NSW study, the absence of an auditable longitudinal rehabilitation treatment rendered treatment-dependent inference nonoperational rather than forcing inference through a proxy exposure.
Dier Tang, Guangyue Han
Wed 30 Sep
cs.ITcs.LGmath.ITmath.OCmath.STstat.MLstat.TH
pdf html
- favorites near “PAC-Bayes Compression Bounds So Tight That They Can Explain Generalization” 0.84
- reading theory, chernoff, attain, equation
- interest Kernels, operators and functional data 0.76
Abstract
The classical information bottleneck (IB) measures the relevance of a representation $U$ of $X$ to a target $Y$ by $I(U;Y)$, which does not directly characterize the error of downstream decisions. For a binary hypothesis $Y$ inferred from many separately encoded observations, the optimal error exponent is the Chernoff information between the two conditional distributions of $U$ given $Y$. We study the mutual information constrained Chernoff bottleneck, which seeks an encoder that maximizes this Chernoff information subject to a rate constraint $I(U;X) \leq R$. We show that its optimal value $C(R)$ increases strictly up to $R = H(V)$, where $V$ merges the symbols of $X$ with equal likelihood ratio, remains at the uncompressed exponent beyond, and, unlike the IB curve, need not be concave. We further show that $k+1$ outputs suffice to attain $C(R)$, where $k$ is the cardinality of $V$. We propose an alternating algorithm that updates the encoder via a generalized Blahut--Arimoto algorithm and the Chernoff parameter $s$ via a nonlinear equation, and prove that its iterates remain feasible, with nondecreasing and convergent Chernoff information. Numerical experiments confirm the theory, and on real topic-detection data from the 20 Newsgroups corpus, compressing each word to only $17\%$ of its entropy retains $90\%$ of the error exponent and nearly the accuracy of the uncompressed classifier.
David Yallup, Will Handley
Fri 25 Sep
cs.LGstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.76
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.76
- reading mcmc, density estimation, advances, interacting
Abstract
Sampling from Boltzmann distributions of molecular systems is an inference problem that has seen significant recent developments fuelled by advances in neural density estimation. We develop a novel sampling algorithm, Neural Transport Nested Sampling (NTNS), which combines the classical strengths of nested sampling with modern neural flow-based methods. NTNS uses a flow matching velocity as the drift in a Metropolis--Hastings corrected Langevin kernel inside a nested sampling outer loop, requiring only evaluations of the target energy function and providing scalable estimation of the full partition function of high-dimensional particle systems. We benchmark NTNS on challenging molecular sampling benchmarks, scaling up to Lennard--Jones clusters of 55 interacting particles, where it reduces both interatomic distance and energy Wasserstein errors to reference MCMC by over an order of magnitude relative to the strongest neural baselines at lower wall-clock cost. To our knowledge, NTNS is also the first neural sampler to return a calibrated, temperature resolved partition function estimate at this scale, recovering the phase structure across temperature from a single run.
Recep Özkan, Christian Hirsch
Tue 29 Sep
stat.MLcs.LGmath.PR
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.80
- interest Kernels, operators and functional data 0.78
- reading formula, kink, higher, makes
Abstract
We study how the affine geometry of finite ReLU networks is created at random initialization and reorganized by supervised training. We call a sign-changing zero of a hidden preactivation an activation switch and a point where the scalar network output is nondifferentiable a scalar kink. For one-dimensional input, conditioning on the preceding layers makes each preactivation Gaussian and affine on the cells of a random finite partition. This yields an exact finite-width conditional Kac-Rice formula for the expected number of activation switches along an input interval. For fixed depth and proportionally growing widths, the resulting switch intensities converge to explicit deterministic limits. A visibility estimate shows that the expected number of switches that do not produce scalar kinks is negligible, yielding an explicit leading formula for the expected number of scalar kinks and hence affine regions. In higher input dimensions d >= 2, the analogous conditional surface formula yields the leading expected (d-1)-dimensional Hausdorff measure of the scalar kink set. On the Breast Cancer Wisconsin data, the initialization formula accurately predicts switch counts along held-out segments. After training, switch counts decrease along within-class segments and increase along between-class segments. Thus, training redistributes rather than merely contracts affine complexity.
Malte Londschien
Thu 24 Sep
stat.MEcs.LG
pdf html
- interest Adaptive data collection and anytime-valid inference 0.81
- your work near “Valid Inference with Imperfect Synthetic Data” 0.77
- favorites near “Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models” 0.80
Abstract
We study prediction in new environments when training data contain rare, large shifts. Anchor regression penalizes the average of the squared mean residual across environments. It protects against shifts in an ellipsoid determined by the second moment of the training shifts. Covering rare shifts may therefore require a large penalty, expanding the ellipsoid in every direction and reducing accuracy on common environments. We propose CVaR anchor regression, which replaces the average of the squared mean residuals with a tail average. Unlike CVaR or GroupDRO applied directly to prediction risks, it does not give environments more weight solely because their noise levels are high. We prove an exact worst-case risk guarantee under a linear structural model that allows for heteroscedastic noise. For discrete environments, decreasing the CVaR tail fraction expands the robustness set from an ellipsoid to a scaled convex hull of the training shifts and their negatives. A separate parameter controls its scale. Examples show how the method can improve protection against rare shifts while retaining accuracy on common environments. We illustrate the method on New York City taxi data.
Divan Aristo Burger, Emmanuel Lesaffre, Reynaldo Martina
Tue 29 Sep
stat.ME
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.80
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.79
- reading log, little, history, treatment
Abstract
Reference-based multiple imputation is used in longitudinal clinical trials to assess sensitivity to assumptions about outcomes unobserved after intercurrent events. Most existing methods target continuous outcomes and use multivariate normal working models. Scheduled count outcomes require a model that preserves integer support, skewness, overdispersion, longitudinal dependence, and an exposure-based rate interpretation. We propose visit-specific negative binomial (NB) margins linked by a Gaussian copula. Covariate-adjusted log rate models define the margins, while the copula captures longitudinal dependence. For missing outcomes in the active arm after a prespecified event, termed the trigger, assigned arm continuation retains the fitted active arm marginal mean, jump to reference uses the corresponding reference arm marginal mean, and the intermediate rule interpolates between these means on the log rate scale. Combined with the copula, the resulting NB margins determine an imputation distribution conditional on the retained history. We account for discreteness through randomized probability integral transform augmentation and fit the model with a custom Metropolis-within-Gibbs sampler. In targeted simulations, the model recovered the generating marginal and dependence parameters with little bias. In simulations with incomplete data, treatment effect estimates were closest to the corresponding complete data estimates when the imputation rule matched the mechanism governing outcomes after the trigger. We illustrate the method using repeated incontinence episode counts from a published trial in overactive bladder. Estimated rate ratios comparing active treatment with placebo remained below 1 under all reference-based assumptions, with modest attenuation toward the null and the greatest separation between rules at the final visit.
Linzhe Zhang, Changming Xu
Fri 25 Sep
stat.MLcs.AIcs.LG
pdf html
- reading state, nominal, govern, interventions
- your work near “Timing as an Action: Learning When to Observe and Act” 0.75
- interest Adaptive data collection and anytime-valid inference 0.79
Abstract
Before acting upon an unobservable physical system, an autonomous agent must determine which latent distinctions govern downstream tasks, how many active interventions are necessary to certify them, and when to abstain to prevent catastrophic errors. Classical observability treats state reconstruction as an unconditioned binary predicate, failing when passive observations cannot break latent degeneracies without perturbation, full microscopic inversion is prohibitively costly, and distinguishing task-irrelevant degrees of freedom wastes interaction budgets. We formalize task-conditioned active observability complexity: the minimum worst-case expected interaction cost required to identify task-relevant states under certified error and safe abstention guarantees. We prove that task-predictive equivalence induces the unique minimal sufficient quotient $\mathcal{H}/\!\sim_\tau$, leaving active observability complexity strictly invariant while eliminating superfluous distinctions. In deterministic regimes, this complexity is characterized by an optimal adaptive distinguishing tree and Bellman recursion; in noisy regimes, it obeys a stopped-transcript relative-entropy lower bound and adaptive martingale certificates that compose without independence assumptions. We instantiate a prospective certified observer with staged recovery: a nominal verifier defers candidate compilation, triggering active probing only upon evidence, while a history-measurable score shell prunes hypotheses without sacrificing risk bounds. Stress audits across high-dimensional physical systems and thousands of operational trials demonstrate certified state recovery with zero false acceptances and substantial reductions in sensor reads and model steps.
Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao
Fri 25 Sep
stat.MLcs.LG
pdf html
- reading pretraining, functions, theory, generalization
- interest Language models as statistical objects 0.78
- your work near “Timing as an Action: Learning When to Observe and Act” 0.74
Abstract
Transformers achieve remarkable performance by jointly learning broad families of tasks during pretraining and adapting to unseen tasks from only a short prompt. Yet a rigorous mathematical and statistical understanding of this phenomenon remains limited. This paper aims to study how Transformers exploit shared cross-task structure and how this structure affects the sample complexity of in-context learning (ICL). Specifically, we characterize task-space complexity through covering numbers under a prescribed metric, thereby quantifying the low-dimensional cross-task structure without requiring an explicit parametric representation. The resulting cover provides a set of anchor functions, which we use to introduce a task-identification-and-evaluation procedure: context observations localize an unseen task among the anchor functions, and the response at a query is predicted by aggregating the corresponding anchor function query evaluations. For approximation, we explicitly construct a Transformer with Softmax attention to approximate this procedure. For generalization, we derive an error bound that separates the effects of the number of pretraining tasks and the prompt length. The scaling with respect to the number of pretraining tasks is governed by the intrinsic dimensions of the task space and input domain; once sufficiently many tasks are available, the dependence on the prompt context length becomes dimension-free. To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL. Our theory provides a quantitative explanation of how joint pretraining across related tasks improves in-context generalization.
Tengfei Shao
Wed 30 Sep
cs.SIstat.ME
pdf html
- reading shock, degree, rise, fewer
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.70
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.74
Abstract
Japan's 2020 tourism shock coincided with 27% fewer buyers and 35% lower gross merchandise value at one second-hand luxury resale intermediary. We compared 2019 and 2020 co-purchase projections of recorded brand labels with a Holm-corrected battery of five label-invariant metrics and Portrait Divergence against two permutation nulls rebuilding both graphs. Neither test separated the full years (smallest Holm p 0.407; 0.097 under the design-matched paired-swap null). For one random two-block split of brands, the battery detected a planted redirection of 30% of multi-brand baskets (178 of 10,399 second-group baskets, 1.7%) in 83 of 100 simulations (83-92 of 100 across 11 splits, all at least 80); Portrait Divergence reached 80% power only at full redirection. Observed falls in mean degree and clustering resembled those of 20-30% redirection, but the planted rise in degree assortativity was absent. A composition-only model of the observed attrition would be detected in only 15 of 100 simulations. Results moved with construction: without residual and material labels (post hoc), mean degree separated the years under the paired-swap null (Holm p = 0.030), and Portrait Divergence fell below 0.05, unadjusted, at the widest construction and within April-December. Shock-period network comparisons need rebuilt-graph nulls and a per-metric detection floor.
Xin-Yu Hu, Shuang Liang, Cheng Feng, Shao-Qun Zhang
Fri 25 Sep
cs.LGcs.AIstat.ML
pdf html
- reading forecasting, series, empirical, forecasting performance
- interest Macroeconometrics and time series 0.78
- favorites near “Eight Centuries of Global Real Interest Rates, R-G, and the 'Suprasecular' Decline, 1311-2018” 0.69
Abstract
The recent emergence of Time Series Foundation Models (TSFMs) has significantly advanced multi-step forecasting performance, enabling accurate predictions over extended future horizons. However, existing TSFMs often suffer from significantly inherent uncertainty, which typically manifests as derived forecast branches emerging at each time step and spreading to subsequent steps; different forecast branches often exhibit varying forecasting performance, thereby undermining the credibility of TSFM forecasts. In this paper, we propose the Slicing-Graphing-Alignment (SGA) method to quantify the uncertainty of multi-step TSFM forecasts. The proposed SGA first characterizes the topology of all potential forecast branches using a directed acyclic graph, such that the graph complexity bounds the uncertainty of multi-step forecasts, and then precisely measures the graph complexity by integrating both topological information and TSFM-inherent stochasticity. Experimental results conducted on 11 TSFMs and 27 datasets demonstrate that (i) SGA achieves the best performance when ranking predictive errors with uncertainty estimates; (ii) SGA works with a more extensive and more precise sampling coverage than those of existing UQ methods, deriving a quantification mechanism fundamentally different from those of established ones; and (iii) larger model scales of TSFMs correlate with lower uncertainty estimates of multi-step forecasts, suggesting another empirical scaling law for uncertainty quantification of multi-step TSFM forecasts.
Arnaud Coatanhay, Angélique Drémeau
Fri 25 Sep
math.STmath.DGstat.TH
pdf html
- reading regular, theory, nuisance, local asymptotic
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.74
- interest Kernels, operators and functional data 0.76
Abstract
A regular observed statistical model may converge to a limit in which a previously identifiable signed parameter becomes identifiable only modulo a reflection. We study the local information geometry of this transition. For a twice differentiable Hellinger embedding with an exact limiting reflection, the observed displacement is forced into the two-jet form \(\varepsilon\lambda J_-+\lambda^2J_+/2\), up to higher-order terms. The mixed jet restores the sign away from the symmetric face, whereas the even jet is the first tangent inherited by the quotient. After nuisance elimination, a positive Gram determinant yields a nondegenerate cross-cap two-jet. The associated local asymptotic theory has three regimes governed by \(\tau_n=\sqrt n\,\varepsilon_n^2\): regular signed LAN, a critical curved Gaussian subexperiment, and a quotient regime with the \(n^{-1/4}\) signed scale. We prove that the same parabolic critical experiment persists for predictive likelihoods along a single stationary dependent trajectory. The limiting quotient has a regular Fisher metric in the invariant coordinate, while its pullback degenerates in the signed coordinate. For a solvable CIR--OU benchmark motivated by coherent sea-clutter observations, we derive the quotient Fisher metric and curvature explicitly and show that the curvature is strictly negative. We also determine the restricted holonomy of the full Amari family: \(\operatorname{Hol}_0(\nabla^{(a)})=SO(2)\) for \(a=0\), whereas \(\operatorname{Hol}_0(\nabla^{(a)})=GL^+(2,\mathbb R)\) for \(a\neq0\). The results separate the intrinsic geometry of the limiting quotient from the transverse geometry of its weak unfolding.
Xiaofei Wu
Thu 24 Sep
stat.CO
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.76
- your work near “Automated Solution of Heterogeneous Agent Models” 0.78
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.78
Abstract
This note revisits the convergence of the Parikh--Boyd two-way block-splitting algorithm for large-scale distributed learning through the He--Yuan prediction--correction framework. Simultaneous row--column partitioning is also relevant to hybrid federated learning, where data may be heterogeneous in both samples and features. We lift the reduced iteration to an equal-dimensional product space and reconstruct the primal and dual coordinates omitted by its implementation. The induced orthogonal-complement structure establishes exact iteration-by-iteration equivalence with the published updates. A mixed variational-inequality representation then yields a fundamental descent inequality, global convergence, an ergodic complexity bound, and current-iterate residual estimates under standard convexity, solvability, exact-subproblem, and invariant-initialization assumptions. The analysis also shows that the reduced state recursion is a metric proximal point iteration. No strong convexity, differentiability, or full-rank condition is imposed. The derivation clarifies which algebraic initialization conditions allow the reduced implementation to inherit the full-space convergence and complexity guarantees without modifying its local updates.
Honam Wong, Surbhi Goel, Enric Boix-Adserà
Wed 30 Sep
cs.LGcs.AIcs.CLstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.79
- interest Language models as statistical objects 0.81
- your work near “Automated Solution of Heterogeneous Agent Models” 0.73
Abstract
Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $\Theta(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of $\Theta(h/r)$ more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.
Futoshi Futami, Jerry Huang, Ichiro Takeuchi
Wed 30 Sep
stat.MLcs.LG
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.77
- favorites near “Cross-Prediction-Powered Inference” 0.88
- interest Online learning and statistical learning theory 0.79
Abstract
In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a way to guarantee downstream decisions while controlling computational and sample requirements. At the same time, some multiclass problems have meaningful label hierarchies that play important roles in medicine and image classification, yet how UC evaluates utility within a hierarchy remains insufficiently understood. We show that the difference between realized utility and predicted mean utility admits an exact decomposition into a sum of contributions from the internal nodes of the label tree. This decomposition shows that positive and negative contributions from different nodes can cancel, and that even when UC is small, the utility errors remaining in parts of the hierarchy need not be small. To address this problem, we propose Hierarchical Utility Calibration (HUC), which evaluates each node contribution before summation while retaining the same target utility, subgroup, and predicted-utility interval. We further provide finite-sample evaluation over all predicted-utility intervals and propose HUC-Boost, which updates only violated internal nodes, with theoretical guarantees for both.
Zahra Naji, Montserrat Fuentes, Liangsuo Ma, Hossein Moradi Rekabdarkolaee
Thu 24 Sep
stat.MEstat.CO
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.80
- interest Causal inference and identification in econometrics 0.78
- favorites near “Hypoelliptic diffusion maps I: tangent bundles” 0.80
Abstract
Multidimensional array data, or tensors, arise naturally in neuroimaging and other high-dimensional applications. We propose a parsimonious Bayesian tensor regression model for studies in which a brain image is the response and predictors are vector-valued covariates. The method extends Bayesian envelope dimension reduction to tensor responses, identifying material subspaces that contain regression information while removing variation that is immaterial to the predictors. This formulation leads naturally to a Tucker tensor decomposition and allows spatial dependence and multiple sources of uncertainty to be modeled jointly. We develop a computationally feasible Markov chain Monte Carlo algorithm based on Gibbs sampling and establish posterior consistency for the proposed model. Simulation studies demonstrate substantial gains in estimation accuracy and uncertainty quantification when meaningful dimension reduction is present. We apply the method to Human Connectome Project neuroimaging data to investigate associations between alcohol use and brain activity. The results illustrate the value of Bayesian tensor envelope regression for inference with high-dimensional, spatially dependent imaging responses.
Shahid Khan, Andriy Olenko
Wed 30 Sep
stat.MEstat.AP
pdf html
- reading appear, procedures, considers, different data
- your work near “Valid Inference with Imperfect Synthetic Data” 0.75
- interest Macroeconometrics and time series 0.76
Abstract
Spherical data appear in various applications, including cosmology and the Earth sciences. A standard statistical model for such data employs isotropic spherical random fields. The isotropy assumption may be unrealistic for many real-world datasets. Also, in various applications such as cosmology, only a single realisation of the field is observed, making isotropy impossible to assess through repeated sampling. This paper considers an alternative anisotropic model, referred to as diagonal anisotropy. To test isotropy against diagonal anisotropy, we develop score and cumulative-sum-type tests. Their asymptotic properties and Monte Carlo-based alternatives suitable for moderate sample sizes are established. The performance of the proposed methods is illustrated by numerical studies via applications to simulated data and to actual Planck cosmic microwave background radiation observations. The tests are also applied to three Planck SMICA cosmic microwave background radiation maps and demonstrate how results can vary for different data releases and map construction procedures.
Joseph Uririoghene Obukofe, Anthony O'Hare, Chioma Sandra Dike
Tue 29 Sep
cs.LGstat.ML
pdf html
- your work near “Automated Solution of Heterogeneous Agent Models” 0.78
- interest Online learning and statistical learning theory 0.81
- reading variational, history, inference, sequential
Abstract
We present a three-tier progression of Hidden Markov Models: maximum-likelihood (Baum-Welch), variational Bayesian (VBEM), and a neural variational extension (Neural VBEM), that model each customer's transaction history as a trajectory through a small number of latent behavioural regimes, one of which is empirically identified as fraud-associated. The Neural VBEM HMM replaces the fixed Gaussian-multinomial emission family with a learned encoder, compressing a 741-dimensional transaction representation into a 64-dimensional latent space in which the VBEM HMM's posterior operates; a UMAP projection of this space reveals that the discovered regimes are not discrete clusters but ordered segments of a single continuous behavioural manifold, with confirmed fraud concentrated at its extreme. We show that the model's natural output, that is, the posterior probability of regime membership, is routinely mistaken for a fraud probability, and quantify the resulting miscalibration (the regime-membership interpretation error, MRIE); a corrected posterior-predictive score, closes most of this gap. We further distinguish batch (smoothed) inference, which uses look-ahead unavailable at deployment time, from filtered (forward-only) inference, and report both. On IEEE-CIS transaction data, the neural tier achieves a 14.4$\times$ fraud enrichment in its identified regime; while its AUPRC trails a discriminative XGBoost baseline, we show this gap is structural and not incidental, and argue the model is best positioned as a calibrated triage and interpretability layer rather than a drop-in ranking replacement.
Abhijit Brahme, Alexander Franks
Tue 29 Sep
stat.AP
pdf html
- favorites near “Recent Developments in Machine Learning Methods for Stochastic Control and Games” 0.77
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.70
- reading career, nonparametric, tends, latent-space
Abstract
NBA athlete performance tends to increase through early career as athletes develop and acclimate to the league, followed by decline due to age-related deterioration in athleticism. While this general pattern persists, the precise shape of this trajectory varies by athlete and across different measures of performance. To model performance increase and decline, we introduce the concave process prior, a novel nonparametric prior over concave functions. We then use a latent variable model to characterize dependence in aging profiles across player-metrics, embedding each player in a shared low-dimensional latent space so that players with similar profiles learn similar trajectory shapes, peak ages, and peak values. Posterior analysis of the learned embedding supports latent-space nearest-neighbor retrieval of career-comparable players and informed projections of young players. We apply our model to data across over a dozen performance metrics for over two thousand players in seasons ranging from 1997 to 2026. Our results show that jointly modeling all metrics improves held-out predictive performance over single-metric alternatives, and that the concavity constraint itself improves prediction. We find that athleticism-driven metrics such as blocks and offensive rebounds peak in a player's early twenties, while skill-based shooting metrics peak in the mid-twenties or later.
Gianmarco Callegher, Thomas Kneib
Wed 30 Sep
stat.CO
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.79
- reading variational, inference, approximations, sequential
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.81
Abstract
Structured additive distributional regression flexibly relates all parameters of a conditional response distribution to covariates, but multivariate extensions remain challenging when dependence is complex. We propose a multivariate structured additive distributional regression model based on regular vine copulas. Different vine edges may use different pair-copula families, while every pair-copula parameter may vary with covariates through a structured additive predictor. The model therefore accommodates heterogeneous marginal distributions together with pair-specific asymmetric, tail-dependent, and covariate-dependent dependence structures. For scalable inference, we develop a tree-wise stochastic variational inference procedure based on component-specific Gaussian variational approximations. Marginal models are estimated first, followed by pair-copula regressions sequentially along the vine trees. We also adapt sequential vine selection by fitting candidate covariate-dependent pair-copula regressions and using information criteria based on effective degrees of freedom for both family and tree selection. In simulations, the tree-wise estimator remains close to an oracle using the true recursive conditional inputs, whereas global refinement yields more concentrated approximations and lower frequentist coverage for upstream components. An application to six-dimensional meteorological data from the Netherlands yields a vine combining Gaussian and non-Gaussian pair copulas, with pronounced nonlinear spatial and temporal variation in dependence.
Paul Bach
Fri 25 Sep
stat.MEmath.STstat.COstat.TH
pdf html
- reading basis, slightly, splines, practice
- interest Kernels, operators and functional data 0.77
- favorites near “Monte Carlo Geometry Processing” 0.77
Abstract
Thin plate splines are highly attractive smoothers. However, they have cubic computational cost, which severely limits their use in practice. As a remedy, Wood (2003) suggested thin plate regression splines (TPRS), which provide a low rank approximation. The key step of the TPRS approximation is a truncated eigendecomposition of the radial basis function (RBF) design matrix. However, as Wood (2003) writes, the optimality of the TPRS approximation is a slightly weak one. This is because the RBF coefficients are subject to orthogonality constraints and the TPRS approximation is only optimal if these constraints are ignored. To address this shortcoming, we suggest a slightly different low rank approximation. The suggested approximation is based on a truncated Demmler-Reinsch basis (TDRB), which provides a best low rank approximation of the smoother matrix in terms of Frobenius and spectral norm. We prove that the TDRB smoother achieves the optimal rate of convergence and suggest an efficient algorithm for its construction. This algorithm is based on a truncated Karhunen-Lo\`eve (KL) expansion of the equivalent Bayesian smoothness prior and it has the same computational cost as required for TPRS. We demonstrate the applicabilty of our approach through simulations and a real data example. We find that the performance is very similar to that of TPRS but the suggested approach has some advantages.
Zeyuan Yu, Zhi Chang, Shiwei Lan
Tue 29 Sep
stat.MLcs.LGmath.PR
pdf html
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.76
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.78
- interest Kernels, operators and functional data 0.77
Abstract
Flow matching (FM) has recently emerged as a promising framework for generative modeling due to its conceptual simplicity and strong empirical performance. In FM, samples are transported along a vector field parameterized by a neural network, inducing a probability path that evolves from a simple noise distribution to the target data distribution, governed by an ordinary differential equation (ODE). However, existing FM approaches predominantly rely on probability paths derived from optimal transport (OT) between Gaussian distributions, which may be suboptimal for capturing complex data with inhomogeneous structures such as heavy tail or sharp contrast. In this work, we generalize FM to the broader class of location-scale families for handling data inhomogeneity and introduce a novel class of probability paths defined as geodesics on the manifold of probability distributions. We name this approach probabilistic geodesic flow matching to distinguish it from prior geodesic (Riemannian) FM methods defined in input space. We argue that Euclidean OT-based paths are not necessarily optimal in probability space and may limit modeling flexibility. Through synthetic benchmarks and scientific datasets at different scales, we demonstrate that the proposed method more effectively captures complex distributions, leading to improved or comparable performance compared with SOTA geometry-motivated generative models.
Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti
Mon 28 Sep
cs.AIstat.ME
pdf html
- reading language, mechanics, input output, gives
- favorites near “Data Science at the Singularity” 0.81
- interest Computational methods for heterogeneous agent and dynamic economic models 0.70
Abstract
Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic concept has one valid syntax, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine executes. Compilation changes representation, not meaning: it does not repair invalid definitions or inject hidden defaults. Programs use fixed schemas of named input and output fields, the leaf output fields are the scoring dimensions, and the engine exposes one universal runtime contract --- a named-field input object in, a named-field output object out --- to which external AI-systems adapt at the boundary, so integration mechanics never propagate into benchmark semantics. This paper gives the semantic object model, the ontology and task-to-program validation rule, the scoring and failure semantics, the compilation and execution architecture, and the scope of the current language. An appendix fixes the normative engineering contract for the first engine implementation.
Preben Johnsen Bentdal, Nello Blaser, Xue-Cheng Tai
Tue 29 Sep
cs.LGstat.ML
pdf html
- favorites near “Hierarchies of Relaxations for Online Prediction Problems with Evolving Constraints” 0.80
- your work near “Valid Inference with Imperfect Synthetic Data” 0.75
- interest Language models as statistical objects 0.77
Abstract
We study post-hoc refinement of frozen node classifiers: given only the graph $G$ and class distributions $Q$ predicted by a frozen model, can we improve accuracy without access to node features, model parameters, or gradients? APPNP answers this by propagating logits with a restart towards the initial predictions, minimizing the anchored Dirichlet energy. Instead, we consider the Potts energy, and decompose it into a Dirichlet term, which penalizes disagreement between neighbouring nodes, and a Gini term, which penalizes indecision within each node. This decomposition motivates Propagate, Then Sharpen (PtS), which alternates between propagation of class probabilities and node-wise, mass-preserving sharpening, with only one additional hyperparameter selected using labelled validation nodes. Across nine homophilic graphs, with a frozen MLP backbone, PtS improves mean test accuracy over independently tuned APPNP by $1.71$ percentage points on clean inputs and $3.90$ under severe Gaussian feature corruption. Gains over APPNP become smaller, but remain positive with frozen GCN and GraphSAGE backbones. Sharpening also removes most of the accuracy loss of deep propagation: on clean inputs without restart, accuracy falls by $2.2$ points between $2$ and $100$ propagation steps under PtS, compared with $33.8$ for APPNP.
Jianyu Xu, Smriti Jha, Aarti Singh, Bryan Wilder
Wed 30 Sep
cs.LGstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.81
- reading policy, online, aggregate, local asymptotic
Abstract
Modern learning systems often acquire supervision at multiple resolutions, trading annotation cost against information content. We study cost-aware two-resolution learning, where expensive fine labels reveal a vector response and cheaper coarse labels reveal a scalar aggregate formed with unknown weights, while the target remains the full response. The challenge is that unknown aggregation changes which directions coarse data can identify, so the value of coarse supervision depends jointly on cost, noise, and identification. We characterize this information geometry and develop an estimate-and-track policy that learns the aggregation rule and tracks the optimal resolution mix. We derive a closed-form break-even condition for coarse supervision and prove that the online policy attains the optimal leading cumulative-risk coefficient, with a matching local asymptotic minimax lower bound. Synthetic experiments support the predicted all-fine/mixed transition, show the online learner approaching the oracle-share benchmark, and demonstrate a finite-budget gain over all-fine acquisition when coarse supervision is sufficiently favorable. Our results provide a principled way to balance information and annotation cost across supervision resolutions.
Runbing Zheng, Dmitriy Kunisky
Wed 30 Sep
stat.ME
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.81
- interest Empirical Bayes, shrinkage and decision theory 0.78
- reading sequential, making, matrices, semidefinite
Abstract
Multi-source data integration problems over datasets from different sources covering different but possibly overlapping sets of entities have become increasingly important in many real-world areas, including genomics, single-cell analysis, and healthcare research. In such problems, one often first learns a low-dimensional representation of the entities within each source and then integrates these representations across sources. As the representations from different sources are only identifiable up to some transformation, how to align them across sources using the sources' overlapping entities becomes a key challenge. Existing methods align the sources in a sequential or tree-structured manner, and are therefore sensitive to the chosen order and exploit only part of the available overlapping information. Motivated by this limitation, we propose Global Synchronized Multiple Matrix Integration (GSMMI), which formulates this alignment problem as a global synchronization problem and jointly aligns all sources using all pairwise overlaps at once, thereby making full use of all overlapping information across the sources. We develop an efficient iterative algorithm for GSMMI that is fast and scalable to the large-scale data arising in these applications. We show both theoretically and empirically that GSMMI improves alignment accuracy, with clear improvements even under modest overlap structure. Moreover, we develop GSMMI to be broadly applicable across data types, covering symmetric positive semidefinite, symmetric indefinite, and asymmetric or rectangular matrices, and even settings where sources overlap only in their rows or only in their columns, making it suitable for a wide variety of application scenarios.
Xinyuan Xie
Fri 25 Sep
math.PRcs.DMcs.NEmath.COstat.ML
pdf html
- reading functions, retrieval, memories, questions
- interest Kernels, operators and functional data 0.75
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.67
Abstract
How much information can a single neuron remember? How many memories can neural networks retrieve without creating false memories? These questions are related to a basic question: how many Boolean threshold functions $f(x)=\operatorname{sgn}(a_0+\langle a,x\rangle)$, $x\in\{-1,1\}^n$, are there? In this paper, we show that the number $T_n$ of distinct Boolean threshold functions is \[ T_n=2\binom{2^n-1}{n}\bigl(1+O(n^{-99})\bigr). \] Equivalently, the capacity of a single threshold neuron is $n^2-\log_2(n!)+1+O(n^{-99})$ bits, improving the $O(n)$ error term in the result of Kahn--Koml\'os--Szemer\'edi to $O(n^{-99})$. To prove this, we show that, for $1\le r\le n-1$, and $v_1,\ldots,v_r$ are chosen at random from $\{-1,1\}^n$, \[ \mathbb P\!\left\{ \langle v_1,\ldots,v_r\rangle\cap\{-1,1\}^n =\{\pm v_1,\ldots,\pm v_r\} \right\} =1-O(n^{-99}). \] In the context of the Kanter--Sompolinsky Hamiltonian for memory retrieval, this identifies $r=n-1$ as a sharp threshold, at which, for almost every collection of $r$ memories, the only ground states are these memories and their negatives, confirming a weaker form of the Kalai--Linial--Odlyzko conjecture. It also settles a recent open problem posed by M. Anthony on the specification number of Boolean threshold functions. In addition, we show that, for every $1\le r\le n-1$, \[ \mathbb P\{v_1,\ldots,v_r\text{ are linearly dependent}\} =2\binom r2\,2^{-n}+O\!\left(2^{-n}e^{-cn}\right), \] confirming a conjecture of Kahn--Koml\'os--Szemer\'edi.
Luis A. Barboza, Carlos Pasquier, Baltazar Espinoza, Fabio Sanchez
Thu 24 Sep
stat.AP
pdf html
- reading aggregate, structural, partially identified, did
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.72
- interest Causal inference and identification in econometrics 0.73
Abstract
Epidemic incidence reflects both transmission dynamics and adaptive human behavior, yet these mechanisms may be difficult to distinguish from aggregate case data alone. We calibrated four susceptible--infected--recovered (SIR) specifications to weekly confirmed COVID-19 incidence in New York City from June to December 2020, comparing a single continuous SIR trajectory, a wave-initialized SIR model, and two adaptive-behavior models with either shared or wave-specific transmission. Inference was performed using rejection Approximate Bayesian Computation (ABC), and in-sample reconstruction was assessed using root mean squared error (RMSE) and the weighted interval score (WIS). Reinitializing the epidemic state by wave produced the largest structural improvement over the continuous SIR trajectory, reducing mean-based RMSE by 48.6\% and WIS by 17.6\%. Adding delayed prevalence-dependent behavioral adaptation with shared transmission further reduced mean-based RMSE by 23.4\%, but yielded essentially unchanged WIS relative to the wave-initialized SIR model. Allowing transmission to vary by wave did not provide a consistent additional advantage and produced strongly asymmetric posterior-simulation trajectories. Behavioral sensitivity, response midpoint, and delay remained only weakly to partially identified. The clearest posterior structure was a negative association between transmission intensity and the behavioral midpoint, indicating that higher transmission could be compensated by behavioral responses activated at lower prevalence. Sensitivity to ordered behavioral priors further showed that reconstruction and behavioral inference depend materially on structural prior assumptions. These results suggest that adaptive mechanisms can improve multi-wave incidence reconstruction, while aggregate incidence alone is insufficient to sharply separate transmission from behavioral adaptation.
Mauricio Herrera-Marín
Thu 24 Sep
stat.APphysics.ao-phphysics.data-an
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.73
- interest Macroeconometrics and time series 0.77
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.68
Abstract
Threshold-derived environmental indices convert native-resolution variables into event indicators and aggregated counts, but surrogate tests may impose their null constraints either before or after this transformation. We show that these choices define different inferential targets. A threshold--copula representation separates temporal dependence from the seasonal event-probability vector and makes seasonal concentration an explicit experimental coordinate. In a prespecified benchmark of 13,500 short-memory monthly trajectories, an index-resolution marginal-and-spectrum surrogate and a constrained native-resolution surrogate produced strongly asymmetric paired conclusions: 1084 trajectories rejected only at index resolution, compared with 16 only at native resolution. Holding the expected annual event count fixed while concentrating event probability within the seasonal cycle increased index-only discordance from 4.4\% to 16.2\%. In an independent experiment, equalizing monthly event probability reduced the corresponding rate from 15.9\% to 3.9\%; the mitigation persisted when monthly percentiles were estimated from an independent 30-year reference period. Native-resolution inference was calibrated under the benchmark null but had strongly mechanism-dependent sensitivity: the strongest prespecified history-feedback alternative had a 30.7\% detection rate under the fixed synthetic gate, with substantially lower sensitivity to a persistent-regime alternative. A Coupled Model Intercomparison Project Phase 6 Amazon dry-month case study was non-discriminating, illustrating why null placement and design-specific detectability should be reported together. The results provide a practical framework for surrogate inference on threshold-derived hydroclimatic and environmental indices.
William Bekerman, Rebecca E. Hasson, Leah E. Robinson, Dylan S. Small
Tue 29 Sep
stat.AP
pdf html
- reading children, observational, social, statisticians
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.66
- interest Causal inference and identification in econometrics 0.72
Abstract
The wide-ranging benefits of physical activity and sports participation among children and adolescents have been closely examined and are reported to include improvements in physical and mental health, cognitive functioning, and social connectedness. However, it remains largely unknown how these activities may affect flourishing and academic engagement, which are closely tied to long-term success, health, and well-being, and how these patterns evolved before, during, and after the COVID-19 pandemic. In this article, we provide the protocol for an observational study using data from the National Survey of Children's Health to examine these relationships among American adolescents. To strengthen our findings, we will conduct this investigation across three different time periods, allowing us to assess the replicability of our conclusions. We introduce a novel statistical design, called data turnover, to carry out this analysis. Data turnover allows a single group of statisticians and domain experts to work together to assess the strength of evidence gathered across multiple data splits while incorporating both qualitative and quantitative findings from data exploration. We delineate our analysis plan using this new method and conclude with a brief discussion of additional considerations for our study.
Tong Yang, Tao Jiang, Yuejie Chi, Ashok Cutkosky, Lin Xiao
Fri 25 Sep
stat.MLcs.LGmath.OC
pdf html
- reading hilbert space, hilbert, operator, attain
- interest Computational methods for heterogeneous agent and dynamic economic models 0.75
- favorites near “Geometric MCMC for Infinite-Dimensional Inverse Problems” 0.80
Abstract
We analyze a simple stochastic inertial Krasnosel'skii--Mann (iKM) method for finding a fixed point of a nonexpansive operator in a real Hilbert space. Our method is obtained simply by adding two inertial extrapolations to stochastic KM [Bravo and Cominetti, 2024], and it retains one call to a possibly biased stochastic oracle per update and achieves sharp rates in both the stochastic and deterministic regimes. Specifically, with our proposed parameter schedule, we prove the following last-iterate fixed-point residual bound: \[ {O}\!\left(\frac{1}{K} +\frac{\sigma\log K}{\sqrt K} +\frac{B_K\log K}{K}\right), \] where $K$ is the horizon, $\sigma$ is the noise level and $B_K$ is the accumulated root-mean-square bias. When $B_K=O(\sqrt K)$, this yields $\widetilde O(\epsilon^{-2})$ sample complexity that matches, up to a logarithmic factor, the stochastic-oracle lower bound given under the unbiased subclass of our model [Foster et al., 2019, Theorem 2]. It also improves the best-known $O(\epsilon^{-4})$ random-iterate guarantee for stochastic KM [Bravo and Cominetti, 2024, Corollary 5.4]. To our knowledge, this is the first single-loop method for general nonexpansive fixed-point problems to attain this near-optimal sample complexity without variance reduction or batching. When the oracle is exact, the same method attains the worst-case-optimal $O(K^{-1})$ last-iterate residual rate [Park and Ryu, 2022, Theorem 4.6], improving the $O(K^{-1/2})$ rate of classical KM [Cominetti et al., 2014; Bravo and Cominetti, 2018].
Haoyuan Chen, Alexandre Thiéry
Thu 24 Sep
stat.MLcs.LGmath.DSstat.ME
pdf html
- interest Macroeconometrics and time series 0.79
- your work near “Differentiable State Space Models and Hamiltonian Monte Carlo Estimation” 0.75
- favorites near “Spectral State Space Models” 0.75
Abstract
Data assimilation estimates a dynamical state from partial and noisy observations. Classical ensemble filters are efficient but restrict analysis updates through finite sample covariance and affine Gaussian distribution. We introduce the Flow Ensemble Filter (FlowEF), which uses conditional flow matching to transport the forecast ensemble from a classical baseline filter to an analysis ensemble. FlowEF uses a localized Gaussian source during training, transports forecast ensemble members from a baseline filter at deployment, and conditions its velocity field on ensembles from that baseline filter and the observation. The proposed model therefore learns a nonlinear update while mapping each baseline ensemble independently. For sparsely observed dynamical systems, FlowEF improves both deterministic and probabilistic metrics over all four classical ensemble filters. It also achieves the best performance among the state-of-the-art generative data assimilation models.
Max Cairney-Leeming, Simone Bombari, Marco Mondelli
Wed 30 Sep
cs.LGcs.AIstat.ML
pdf html
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.78
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.80
- interest Online learning and statistical learning theory 0.79
Abstract
Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $\rho$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $\rho \asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $\rho \ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $\rho \gg d$. Remarkably, the transition moves to $\rho \asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).
Siyi Wang, Alexandre Leblanc, Paul D. McNicholas
Tue 29 Sep
math.STcs.CGstat.TH
pdf html
- reading curves, functional data, functional, boundedness
- interest Kernels, operators and functional data 0.75
- your work near “Computational Methods for Economic Models with Function Valued States” 0.72
Abstract
We introduce \emph{curve band depth} (CBD), a band-based data depth for samples of \emph{unparameterized} planar curves. CBD is motivated by band depth and modified band depth for functional data, but targets trajectory data. Unlike the halfspace-based curve depth of \citet{de2021depth} and the curve stabbing depth of \citet{durocher2023csd}, CBD is defined through a geometric band region generated by two curves, and measures the arc-length proportion of a target curve lying inside such bands. We develop a CBD family consisting of an integral version (int-CBD), an infimal version (inf-CBD), and a fast-walk variant (FW-CBD). The fast-walk band is a narrower band construction contained in the global convex-combination band. We establish boundedness, vanishing at infinity, and similarity invariance for these constructions, together with a Borel-measurability result for the induced depth maps under a mild measurability assumption. A length-penalized variant is proposed for samples with heterogeneous curve lengths. We implement the methods via arc-length sampling and polygonal approximations, and evaluate them through classification of overlapping handwriting data and MNIST-derived digit curves, online-signature screening on \texttt{MOBISIG}, and an exploratory clustering task based on decomposed band contributions.
Ronald Schneider
Tue 29 Sep
stat.AP
pdf html
- reading structural, sequential, inference, measured
- interest Bayesian computation and structural estimation 0.78
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.75
Abstract
Structural health monitoring of deteriorating structures requires the latent deterioration process to be inferred indirectly from measured structural responses and propagated to quantities of engineering interest. This paper presents a probabilistic framework that formalizes this process by distinguishing Bayesian inference of uncertain deterioration and model parameters from diagnosis of the current deterioration state, prognosis of its future evolution, and prediction of structural reliability. The deterioration model thereby provides the common probabilistic link between monitoring-based inference and these downstream tasks. To enable sequential inference as monitoring information accumulates, Bayesian Updating with Structural Reliability Methods (BUS) is extended through a nested formulation of the BUS events associated with successively growing monitoring datasets. The resulting Sequential BUS method allows Subset Simulation to continue from the augmented conditional sample population obtained at the preceding monitoring stage rather than restarting from the prior distribution. The posterior population can then be propagated to monitoring-informed reliability analysis. The framework is demonstrated for fatigue deterioration of a redundant jacket-type structure, where modal properties obtained from simulated vibration monitoring are used to infer a high-dimensional deterioration model of welded connections. The results show how successive monitoring observations constrain plausible deterioration histories, provide probabilistic system-level diagnosis, and update crack-growth prognosis and structural reliability while retaining ambiguity between deterioration scenarios that produce similar structural responses. The proposed framework thus provides a coherent probabilistic route from response-based monitoring data to deterioration diagnosis, prognosis, and reliability prediction.
Yue Xing, Pengfei He, Zitao Li
Thu 24 Sep
cs.CRcs.AIcs.LGstat.ML
pdf html
- reading social, recommendation, causes, large
- interest Language models as statistical objects 0.73
- your work near “Valid Inference with Imperfect Synthetic Data” 0.71
Abstract
With recent advancements in large language models (LLMs) and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader access to act on users' behalf on the internet. However, the vulnerability of automated agents deployed on social media platforms (e.g., for managing a user's personal account) remains underexplored. Existing studies on agent poisoning typically assume that the adversary can expose poisoned content to the agent. Although such an attack is direct and effective, it is more easily detected and mitigated. In the context of social media platforms, this leaves open whether the recommendation system itself would surface such content to the agent in a more subtle manner. Through theoretical analysis, we show that the like-score mechanism used in OASIS can be exploited, and we characterize the conditions under which a multi-stage chain of poisoned posts can steer the agent's feed. Based on these insights, we further develop an algorithm that crafts realistic poisoned posts. Experiments support our theoretical findings and demonstrate the effectiveness of the proposed algorithm. Notably, by exploiting the like-score feedback loop, the attack causes the recommendation system to select poisoned posts even when their user-post similarity falls below the retrieval threshold.
Ge Wang
Thu 24 Sep
stat.MLcs.LGcs.NEmath.STstat.TH
pdf html
- your work near “Automated Solution of Heterogeneous Agent Models” 0.78
- reading variational, state, govern, degree
- interest Kernels, operators and functional data 0.75
Abstract
A normal--inverse-gamma (NIG) latent hierarchy has four parameters, but its induced latent law does not identify all four. For $\sigma^2\sim\mathrm{InvGamma}(\alpha,\beta)$, $\mu\mid\sigma^2\sim\mathcal{N}(\gamma,\sigma^2/\nu)$, and $z\mid\mu,\sigma^2\sim\mathcal{N}(\mu,\sigma^2)$, the marginal law of $z$ depends on $(\nu,\beta)$ only through $c=\beta(1+1/\nu)$. Hence the reconstruction-visible parameter space is the three-dimensional quotient $(\gamma,\alpha,c)$, with a one-dimensional fiber degree of freedom. For a fixed hierarchical variational objective, exact partial minimization of the forward KL divergence to a complete NIG prior selects a unique prior-relative representative on each fiber, yielding an exact three-coordinate reduction with the same optimum as the four-coordinate objective. Writing $\rho_0=2\beta_0/\nu_0$ and $T=c/\{\alpha[(\gamma-\gamma_0)^2+\rho_0]\}$, we show that inverse canonical allocation $1/\nu_{\rm can}$ is an explicit strictly increasing function of $T$. For $\alpha>1$, the ratio $u_{\rm epi}/u_{\rm var}=1/\nu_{\rm can}$ is therefore determined by the quotient state and prior; for rank-only use under a common calibration, $T$ contains the same coordinatewise ordinal information. The residual prior gauge is characterized rather than eliminated: $(\gamma_0,\rho_0)$ govern ordinal dependence, while $(\nu_0,\alpha_0)$ determine numerical calibration and the analytic ceiling of $1/\nu_{\rm can}$.
Luca Aiello, Sudipto Banerjee
Mon 28 Sep
stat.MEstat.AP
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.74
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.69
- reading mcmc, south, researchers, inequalities
Abstract
Spatial disease maps help public-health researchers identify geographic inequalities, but standard Bayesian smoothing can obscure localized disparities when neighboring communities have sharply different socioeconomic or behavioral profiles. Analysts therefore need to determine where smoothing should be interrupted and repeat that analysis as maps, adjacency structures, and outcomes change. We develop a covariate-informed Bayesian boundary model and an amortized posterior approximation trained across heterogeneous areal graphs. The model distinguishes local interruptions in smoothing from broader residual spatial dependence; the trained approximation handles maps with different numbers of regions. Simulations examine posterior calibration, boundary-probability recovery, replicated-data behavior, and MCSE-controlled agreement with prior-matched MCMC. In contrast to traditional approaches that analyze these data separately, we demonstrate the effectiveness of using a single trained deep learning network to analyze respiratory hospitalizations in Greater Glasgow; lung cancer incidence in California; and tracheal, bronchial, and lung cancer mortality in South Korea, comprising 58 to 241 regions. Selected boundary density is greatest in Glasgow and lowest in South Korea despite substantial residual spatial dependence in both, showing that local interruption and broader spatial persistence need not vary together. Across all three applications, edge-level boundary probabilities agree substantially with dataset-specific analyses, although posterior spread and thresholded boundary sets differ. These results support reusable Bayesian boundary analysis across the evaluated disease-map class and identify the validation needed before deployment to new applications.
Jose A. Roldan-Nofuentes, Saad-Bouh Regad
Thu 24 Sep
stat.ME
pdf html
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.79
- interest Empirical Bayes, shrinkage and decision theory 0.80
- favorites near “Optimal Decision Rules for Weak GMM” 0.82
Abstract
Comparison of predictive values of diagnostic tests is a topic of interest in Medical Statistics, and has been the subject of different studies. In clinical practice, it is frequent to observe categorical covariates when comparing diagnostic tests. In this framework, a global hypothesis test is proposed to simultaneously compare the predictive values of two diagnostic tests when in all of the individuals categorical covariates are observed. This hypothesis test is solved through regression models and also by weighted least squares method for the analysis of categorical data. Simulation experiments were carried out to study the asymptotic behavior of these methods when a binary covariate is observed and when a covariate with three categories is observed, and these were compared to the behavior of the global test when the covariate is ignored. In general, the method based on regression models has shown to have better asymptotic behavior than the other methods. Furthermore, we studied the application of the method based on the regression models when no covariate is observed, for which the individuals in the sample are randomly assigned to a binary dummy random variable. Simulation experiments carried out showed that this method has greater power than the method without the covariate. The results were applied to two examples.
Yi Zhang, Rui Guo, Mengchu Xu, Zhaofeng Liu, Yonina C. Eldar
Mon 28 Sep
eess.SPcs.LGstat.ML
pdf html
- reading mcmc, computations, nonlinear, posterior inference
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.74
- interest Language models as statistical objects 0.68
Abstract
Split Gibbs sampling enables diffusion posterior inference for general nonlinear inverse problems by decoupling prior and likelihood computations, allowing a pretrained diffusion prior to be reused across measurement models. However, its likelihood update often relies on iterative MCMC, which can hinder parallelization, require algorithm-specific tuning, and incur substantial computational cost. In this work, we propose a learning-based framework to replace this MCMC step by reformulating both Gibbs updates as Gaussian denoising problems and implementing them through ODE diffusion. The prior step reuses a pretrained denoiser, while the likelihood denoiser exploits known likelihood structure through a lightweight deep-unfolded network. Experiments on nonlinear phase retrieval demonstrate the effectiveness of the proposed method as an alternative to MCMC-based split Gibbs at lower likelihood-update cost.
Jinho Cha, Misun Roh
Wed 30 Sep
stat.AP
pdf html
- reading social, burden, adjustment, partly
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.72
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.67
Abstract
Hospice quality should be interpreted in relation to both provider organization and the local conditions under which care is delivered. This study develops a provider-county performance assessment framework by linking national Centers for Medicare & Medicaid Services (CMS) hospice data and Consumer Assessment of Healthcare Providers and Systems (CAHPS) Hospice Survey outcomes with county measures of rurality, social vulnerability, health burden, and workforce and health-resource context. The adjusted analysis includes 2,928 providers in 1,078 counties and combines geographic mapping, blockwise regression, six secondary CAHPS outcomes, and eight sensitivity analyses. Adding county context increased adjusted R-squared from 0.077 to 0.202. After full adjustment, for-profit hospices had overall caregiver ratings 3.832 percentage points lower than nonprofit hospices. This negative association appeared across all six secondary CAHPS domains and remained significant in every sensitivity specification. Higher county social vulnerability was also associated with poorer caregiver experience, although its magnitude depended partly on the specification of community health burden. These findings show that county context materially improves hospice performance assessment but does not eliminate the ownership difference. The framework supports context-aware monitoring, peer comparison, and targeted quality improvement.
Chih-Hsuan Huang, Chih-Wei Chen, Szu-Chi Chung
Wed 30 Sep
cs.LGstat.ML
pdf html
- favorites near “Why Does Deep Learning Work? - A Perspective From Group Theory” 0.79
- interest Kernels, operators and functional data 0.78
- reading manifolds, references, dimensionality, heterogeneity
Abstract
DANCo (Dimensionality from Angle and Norm Concentration) jointly calibrates nearest-neighbor distance and angular statistics and consistently reaches state-of-the-art accuracy on clean intrinsic-dimension (ID) benchmarks. Practical data, however, introduce neighborhood-relative noise and sample-amplitude heterogeneity that can distort these geometric signals. We reformulate DANCo componentwise, retaining separate distance and angular discrepancy curves so that the source of an estimate can be identified and interpreted. For the distance component, we derive a closed-form Kullback-Leibler divergence for the generic-order ratios of the generalized ratios ID estimator (Gride); when both angular parameters are matched (Full), Gride reduces mean percentage error from $27.7\%$ to $17.6\%$ at noise equal to $40\%$ of typical neighbor spacing on 24 manifolds. For the angular component, two sampling regimes motivate aligning mean direction while retaining concentration matching (Profiled). On a Gaussian scale mixture with generating dimension 70 embedded in 100 dimensions, profiling raises the Minimum Neighbor Distance (MiND) estimate from $22.8$ to $66.7$, while removing the known amplitudes restores MiND-Full to $71.9$; the control thus attributes the Full shortfall to amplitude heterogeneity. On CIFAR-10 and ImageNet, amplitude-reducing normalizations move angular location toward the references and narrow the Full-Profiled gap, an observational counterpart to the controlled mixture. Across four pretrained convolutional neural networks, Gride-Profiled, the two-nearest-neighbor estimator (TWO-NN), and the maximum-likelihood estimator (MLE) exhibit similar rise-and-fall profiles, while Full-Profiled differences identify the layers most sensitive to angular calibration.
Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse Şen, Marco Cuturi, Daniel Kuhn
Mon 28 Sep
cs.LGmath.OCstat.ML
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.78
- interest Language models as statistical objects 0.78
- your work near “Timing as an Action: Learning When to Observe and Act” 0.74
Abstract
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be reformulated as an optimization problem over transport maps that push empirical samples to adversarial ones, and we prove that optimal maps are cyclically monotone. We also show that standard adversarial training---based on per-sample local optimization---violates cyclical monotonicity and wastes transport costs unless the adversary is severely restricted. We propose two remedies. First, we introduce multi-start particle ascent, which alternates parallel gradient ascent with reassignment to enforce cyclical monotonicity across samples. Second, we parameterize adversarial maps as gradients of input-convex neural networks, which guarantees cyclical monotonicity by construction. Experiments on robust regression, image classification, and robust control show that our methods consistently outperform standard adversarial training and state-of-the-art baselines, achieving improved robustness and better generalization under distribution shift.
Shoichi Otomo
Tue 29 Sep
cs.CYecon.EM
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.75
- reading macroeconomic, science, higher, little
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.64
Abstract
This manuscript is an English translation and extended version of a paper originally published in Japanese (Otomo, 2024). With the advancement of information and communication technology (ICT) and data analysis techniques, handling massive datasets (big data) has become feasible in spatial information science. Consequently, infrastructure is being established to generate new utility value from spatial data. Furthermore, nationwide initiatives are actively promoting the use of open data - public information provided under terms allowing secondary usage. However, in developing regions and specific municipalities, local statistical data remain scarce, unreliable, or difficult to acquire. Additionally, long-standing concerns exist regarding the inherent limitations of standard macroeconomic indicators when performing cross-regional or international comparisons. Meanwhile, high-frequency satellite data has become widely accessible. Among various observational products, nightlight data has attracted significant attention due to its versatile applications. Generally, nightlight radiance correlates with urbanization. Prior work indicates that areas with larger populations and higher commercial development display higher radiance, whereas public infrastructure - such as sanitation systems and healthcare facilities - shows little correlation with nightlight intensity (Otomo, 2021). Building on these findings, this paper analyzes the relationship between nightlight data and the spatial distribution of sports facilities, the evolution of the fitness industry, and regional environmental disparities in sports access. The analysis confirms that the spatial distribution patterns of sports facilities differ significantly between the public and private sectors.
Bruno Marcondes e Resende, Helton Graziadei, Thiago Rodrigo Ramos, Rafael Izbicki
Wed 30 Sep
stat.MLcs.LG
pdf html
- reading control, conditional risk, curves, gives
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.75
- interest Online learning and statistical learning theory 0.77
Abstract
Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or missed labels in multilabel classification. Conformal risk control (CRC; Angelopoulos et al., arXiv:2208.02814) gives distribution-free guarantees for such losses, but it calibrates a single threshold shared by all inputs. Because conditional risk varies with the input, this marginal guarantee often overprotects easy cases and underprotects hard ones. We propose ReCIRC (Rectified Conformal Risk Control), which inverts each input's estimated local risk curve to reparameterize the calibrated threshold as a risk budget $a$ representing a common target conditional risk, and then applies CRC unchanged to the resulting family. ReCIRC retains CRC's finite-sample marginal guarantee regardless of the accuracy of the estimated curves, while accurate curves yield approximate conditional risk control and, under additional conditions, asymptotically exact conditional risk control; they also support a risk-calibration diagnostic. Across three synthetic and five real-data settings spanning segmentation, multilabel and multiclass classification, and regression, ReCIRC attained the lowest average worst-group risk and mean positive group excess in every setting, while maintaining marginal risk close to the target, whereas changes in prediction size were application-dependent.
Alokesh Manna
Mon 28 Sep
cs.NEstat.ME
pdf html
- interest Kernels, operators and functional data 0.78
- reading lasso, net, pixel, dynamic
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.71
Abstract
Neurons in the visual system are selective for specific spatial and temporal stimulus features, described by their \emph{receptive field}. Estimating one means a coefficient per pixel per time bin from few trials -- a high-dimensional problem requiring regularization. Sparse regularizers such as the LASSO handle the dimension but select pixels independently at each time point, with nothing to keep the region coherent in space or smooth in time; it can fragment or reorganize discontinuously even when the true response evolves smoothly, a failure since this evolving pattern is what a receptive-field estimate should capture. We formulate dynamic receptive-field estimation as a high-dimensional Bayesian problem: a Poisson model combining a Gaussian Markov random field in space with an autoregressive process in time, so the estimated field is smooth and coherent across space and time. On recordings from $155$ salamander retinal ganglion cells, fitting this model independently per neuron recovers a coherent surface, where a pixel-level Poisson-LASSO comparison instead returns a fragmented one. Summarizing each neuron's surface by its space-averaged temporal response and clustering these curves with a model-based functional-clustering procedure, BIC selects three balanced temporal-response phenotypes ($85$, $32$, $38$ neurons), against a degenerate grouping from clustering the raw surfaces. A simulation study with known ground truth confirms the same pattern, with the model beating an unregularized Poisson GLM, LASSO, and the elastic net on recovery and estimation accuracy, though LASSO controls false positives better. The per-neuron field identification, its contrast with LASSO, and the functional-clustering population typing constitute this paper's contribution.
Thibault Pautrel, Florent Bouchard, Ammar Mian, Guillaume Ginolhac
Fri 25 Sep
stat.MLcs.LG
pdf html
- reading personalised, fewer, subject, higher
- favorites near “Spectral State Space Models” 0.73
- interest Language models as statistical objects 0.73
Abstract
Federated learning (FL) lets EEG decoders learn from recordings of several subjects without pooling them. We consider two light EEG decoders, the Riemannian SPDNet and the Euclidean EEGNet. Both split into a trunk, which builds a latent representation, and a head, which classifies it. Inter-subject variability, however, makes a single shared FL model a poor fit for each subject. Personalised FL addresses this: all subjects learn a common trunk, and each subject keeps its own head. We adapt it for SPDNet and study its effects against standard FL and centralised training, with EEGNet as a Euclidean baseline. Experiments cover three motor-imagery datasets that span diverse regimes in channels, subjects and classes. We observe that personalised SPDNet reaches higher accuracy than both standard FL and centralised training, while converging in fewer rounds and communicating fewer parameters than standard FL. It also outperforms every EEGNet configuration on two of the three datasets, although centralised EEGNet outperforms centralised SPDNet.
Yuling Jiao, Wensen Ma, Houduo Qi, Defeng Sun
Fri 25 Sep
cs.CVcs.LGstat.ML
pdf html
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.84
- reading references, pretraining, assign, auxiliary
- interest Language models as statistical objects 0.74
Abstract
Explicit geometric references offer a direct way to structure self-supervised representations. Existing adversarial distribution-matching formulations, however, require costly encoder-critic optimization. We introduce Flow-Based Distribution Matching (FBDM), a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression. An ETF-inspired reference allows its number of components K' to exceed the auxiliary flow dimension d* while retaining structured geometric separation. We assign both augmented views of each image to the same target, while limiting how many images each reference center can receive. An explicit alignment loss further pulls the two views' representations closer together. Experiments across benchmarks ranging from CIFAR to ImageNet show that FBDM achieves performance nearly on par with DM and remains competitive with existing SSL methods. Matched training-cost comparisons show a 1.48- to 1.83-fold speedup over DM with a negligible increase in GPU memory usage. We also provide a theoretical explanation for the usefulness of the learned representations: under stated conditions, we bound the downstream misclassification rate in terms of the FBDM pretraining loss.
Marco Mandap, Jerahmeel Hipolito, Arcel Galvez, Charlie Margaret Balagtas, Michael Joshua Buluran, Jeff Roel Durmiendo, et al.
Mon 28 Sep
stat.MEcs.CL
pdf html
- reading document, control, policy, questions
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.73
- interest Online learning and statistical learning theory 0.74
Abstract
We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal control of any false alert across the permitted inspection path and derive necessary calibration counts for rejection. We also state oracle testing, distribution-shift, and independent-audit bounds with their additional assumptions. Both constructions protect stopping within their specified scope; neither proof constructs an e-process or justifies multiplying conformal ranks. Detection power and computational savings remain questions for empirical evaluation.
Yordan P. Raykov, Max A. Little
Mon 28 Sep
math.STcs.LGstat.TH
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.76
- favorites near “Random Walks on Simplicial Complexes and Harmonics” 0.81
- interest Kernels, operators and functional data 0.72
Abstract
Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential. We use data-derived geometric graphs to precondition an exact subset dynamic program: as a result only connected vertex subsets are admitted as clusters, while sum-of-squared-errors (SSE) loss is unchanged. A remaining-set recurrence minimises fixed-\(K\) or penalised SSE, with exact factorisation over the connected components of each remaining set. The central question we study is how much computational support can be removed while preserving an unrestricted optimum. Graph inclusion gives monotone coverage and support relations, and a bottleneck threshold identifies the first covering graph in a nested hierarchy. For fixed \(K\) and dimension, under compact ball support and density bounds, retaining \(q=O(\log n)\) nearest neighbours per observation preserves an empirical SSE optimum with probability tending to one, using an \(O(\log n/n)\) fraction of complete-graph edges. Truncated Gaussian mixtures with unequal weights and covariances satisfy these conditions. The rate we provide is a sufficient upper bound rather than a result implying polynomial optimisation complexity. Objective-matched synthetic and full-data comparisons assess coverage, compression, and reference-label agreement. As a secondary application, we illustrate how the proposed scaffold preconditioning can be utilized to improve the efficiency of split-merge proposals that preserve unrestricted mixture posteriors.
Akash Kumar
Wed 30 Sep
cs.LGmath.OCstat.ML
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.77
- favorites near “Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs” 0.78
- your work near “Automated Solution of Heterogeneous Agent Models” 0.75
Abstract
Does representation learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates have a step length set by the gradient's rank rather than its norm. Near the edge of stability, full-batch Muon on teacher-student problems enters approximately period-2 loss oscillations that persist for thousands of steps: the cycle-mean loss stays flat or rises, yet the weights keep moving and the learned features continue to align with the teacher subspace. For linear teacher-student learning toys, we derive explicit cycle and alignment formulas and conditional plateau and decay bounds. For a population mean-field ReLU model, we prove that, under stated dimension, initialisation and small-head conditions, the leading eigenspace of the average gradient outer product (AGOP) recovers the teacher subspace exactly during a loss plateau, before the loss later drops. In all 33 ReLU, GELU and SiLU teacher configurations we study, direction-only alignment metrics show the student AGOP aligned with, or still aligning to, the teacher subspace during the period-2 oscillations; projected head refitting on selected configurations shows that the learned directions are useful for prediction, and further measurements distinguish AGOP alignment from weight-mass concentration. In deep residual ReLU students, freezing the downstream layers while the first layer trains with full-batch exact polar updates recreates a nearly flat cycle-mean loss with improving input-AGOP alignment; freezing and unfreezing switch between this plateau and loss decrease, and the effect is sensitive to momentum and to the choice of orthogonaliser.
Zhan Gao, Danyang Huang, Rui Pan, Hansheng Wang
Wed 30 Sep
stat.ME
pdf html
- reading author, theory, data analysis, degree
- your work near “Valid Inference with Imperfect Synthetic Data” 0.76
- interest Semiparametric inference and debiased machine learning 0.78
Abstract
Degree heterogeneity is one of the most important properties of network data. It is widely observed that degree heterogeneity is often related to the nodal features. In this study, we investigate the estimation and statistical inference for a popularity regression-based latent space model with nodal features. Given the complex dependence structure induced by the latent space model, we aim to derive analytically tractable objective functions that effectively account for this structure for sparse networks. Specifically, we propose a total of four estimators. The first two estimators are developed by utilizing only the first-order structure of the network (e.g., the nodal degree), while the last two estimators are developed by leveraging the higher-order network structures (i.e., reciprocity and transitivity). Rigorous asymptotic theory is established based on various non-standard U-statistics. We find that different estimators might have different convergence rates. The extension to higher-order moments-based estimators is also discussed. Extensive numerical experiments and a real data analysis of link prediction for an author citation network are conducted for illustration purposes.
Shruti Nath, Docko Sow, Koomi Toussaint Amoussouvi, Fenwick Cooper, Josiah Kiarie Kimani, John Bagiliko, et al.
Tue 29 Sep
stat.APcs.LGphysics.ao-ph
pdf html
- reading ifs, lead, forecasting, functional
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.67
- interest Macroeconometrics and time series 0.74
Abstract
Artificial intelligence (AI)-based weather prediction is approaching the skill of physical numerical weather prediction (NWP) systems at a fraction of the computational cost. This is particularly promising for Africa, where rainfall extremes are intensifying and many forecasting centres lack the infrastructure to run physical models at extended lead times. We present a calibrated comparison of GraphCast, GenCast and the Functional Generative Network (FGN) against the physical NWP model IFS for rainfall prediction across Africa. Deterministic and probabilistic forecasts are postprocessed using Isotonic Distributional Regression and evaluated with the Continuous Ranked Probability Score against IMERG, RFEv2 and CHIRPS across seasons, wet and dry regimes, elevation zones and lead times. All models retain skill beyond climatology across most seasons and at extended lead times. AI models generally outperform IFS in wet regions, whereas IFS performs better in dry, high-elevation areas, where its finer resolution better represents orographic controls on rainfall. Across observational datasets and seasons, AI models achieve a median improvement of approximately 5% over IFS. GraphCast achieves calibrated skill comparable to the ensemble-based FGN, although FGN provides greater significant skill at longer lead times. These results highlight the potential of calibrated AI weather prediction to provide accessible and computationally efficient rainfall forecasts across Africa, while demonstrating the continuing importance of spatial resolution, ensemble design and regional characteristics.
Fernando Vera-Vera, José A. Roldán-Nofuentes
Thu 24 Sep
stat.ME
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.75
- reading causes, software, matrices, implement using
- interest Kernels, operators and functional data 0.75
Abstract
Biometrical sciences and disease diagnosis in particular, are often concerned with the analysis of associations for cross-classified data, for which distance association models give us a graphical interpretation for non-sparse matrices with a low number of categories. In this framework, usually binary exploratory and response variables are present, with analysis based on individual profiles being of great interest. For saturated models, we show the usual linear relationship for log-linear models is preserved in full dimension for the distance association parameterization. This enables a two-step procedure to facilitate the analysis and the interpretation of associations in terms of unfolding after the overall and main effects are removed. The proposed procedure can deal with cross-classified data for profiles by binary variables, and it is easy to implement using traditional statistical software. For disease diagnosis, the problems of a degenerate solution in the unfolding representation, and that of determining significant differences between the profile locations are addressed. A hypothesis test of independence based on odds ratio is considered. Furthermore, a procedure is proposed to determine the causes of the significance of the test, avoiding the problem of error propagation. The equivalence between a test for equality of odds ratio pairs and the test for equality of location for two profiles in the unfolding representation in the disease diagnosis is shown. The results have been applied to a real example on the diagnosis of coronary disease, relating the odds ratios with performance parameters of the diagnostic test.
Yanzhao Zheng, Yuanqiang Yu, Tianze Xu, Chao Ma, Zhentao Zhang, Jihuai Zhu, et al.
Thu 24 Sep
cs.LGstat.ML
pdf html
- reading counterfactual, aggregate, turn, zero
- your work near “Timing as an Action: Learning When to Observe and Act” 0.74
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.73
Abstract
Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. Existing approaches either augment supervision with sequence- or token-level RL rewards from external verifiers or learned graders, or use on-policy distillation (OPD) against a single full-context teacher whose probability mass becomes diluted as more constraints become simultaneously active. We propose CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation. Rather than enriching the teacher with information beyond what the student sees, CC-OPD ablates each constraint from the teacher's conditioning in turn, and constructs the per-constraint signal from the resulting per-token probability differentials. The resulting per-token leave-one-out log-likelihood shifts are summed, clipped, and added to the vanilla OPD reward as a token-level shaping term. All shaping terms are obtained from the frozen teacher, without an external verifier during distillation, and the reward equals vanilla OPD wherever the aggregate shift is zero. Across two Qwen model pairs and seven benchmarks, CC-OPD achieves the highest average among all evaluated student-training methods. A 1.5B student trained with CC-OPD surpasses its own 7B RL-trained teacher on the MulDimIF benchmark.
Zilan Cheng, Li-Lian Wang, Zhongjian Wang
Tue 29 Sep
cs.LGcs.NEstat.ML
pdf html
- favorites near “Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation” 0.79
- interest Kernels, operators and functional data 0.75
- reading functions, elementary, connection, simpler
Abstract
We study the minimum number of hidden neurons required for arbitrary-accuracy approximation of multivariate H\"older-continuous functions on $[0,1]^d$ and the associated encoding complexity. For $d\geq 2$, we construct a fixed, explicitly defined activation function for which a closed-form network with two hidden layers of widths $d$ and $1$ achieves arbitrary accuracy in the uniform norm. We prove that $d+1$ is the exact minimum total number of hidden neurons among standard feedforward networks with locally integrable activations and affine outputs. We further give a simpler construction using a single elementary activation that combines the floor and exponential functions. This construction requires three hidden layers of widths $d$, $1$, and $2$, only two neurons above the minimum. If a skip connection is allowed, widths $d$, $1$, and $1$ suffice. These constructions use explicit grid addressing and integer encoding of quantized function values. For a bounded $\alpha$-H\"older class, they require $O(\varepsilon^{-d/\alpha}\log(1/\varepsilon))$ bits, matching the metric-entropy lower bound up to a logarithmic factor.
Peida Wu, Xinyang Xiong, Pengcheng Zeng
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Spectral State Space Models” 0.76
- interest Kernels, operators and functional data 0.76
- reading functional, functional data, learning functional, functions
Abstract
Functional data are intrinsically infinite-dimensional, and often exhibit phase variation, where corresponding events occur at different times across observations. Existing linear dimension reduction methods struggle with nonlinear amplitude variation, while functional autoencoders without an explicit warp entangle temporal misalignment with shape. We propose the Amplitude--Phase Functional Autoencoders (AP-FAE), an unsupervised framework for functional data that spans both univariate and multivariate cases, with emphasis on the multivariate setting, and factorizes the latent space into separate amplitude and phase embeddings derived from all channels. A smooth functional decoder reconstructs channel-specific amplitude functions in canonical time, and a shared monotone, endpoint-preserving warp captures phase variation. We prove a bound linking amplitude recovery to registration, reconstruction, and noise errors, and validate it numerically. Across synthetic data and six real-world benchmarks, AP-FAE outperforms state-of-the-art baselines on most clustering and alignment metrics and on all reconstruction metrics. Clustering with amplitude embeddings alone consistently surpasses joint amplitude--phase clustering, confirming the benefit of explicit disentanglement. Code is available at https://anonymous.4open.science/r/APFAE-418C/}{https://anonymous.4open.science/r/APFAE-418C/.
Axel Vuorinen, Antoine Guillon, Emmanuelle Comets, Moreno Ursino
Fri 25 Sep
stat.ME
pdf html
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.85
- your work near “Timing as an Action: Learning When to Observe and Act” 0.73
- interest Reinforcement learning, policy learning and off-policy evaluation 0.71
Abstract
Early-phase dose-finding methods increasingly assess toxicity and efficacy jointly, but comparisons based only on administered dose may inadequately characterize regimens differing in schedule. We developed a Bayesian phase II adaptive platform design for regimen optimization that integrates pharmacokinetic/pharmacodynamic (PK/PD) modelling into toxicity, efficacy, regimen selection and adaptation decisions. The proposed PK/PD-informed Regimen Optimization Platform (PROP) design uses a population PK/PD model to generate patient- and population-level predictions of exposure and biological activity. Acute and cumulative toxicities are analysed using a discrete-time time-to-event model informed by PK exposure. Efficacy is evaluated through Bayesian model averaging of exposure-driven and biomarker-driven time-to-event models. The design supports regimen graduation, discontinuation for futility or safety, and addition of unexplored regimens. Performance was evaluated through simulations motivated by an influenza intensive-care setting. Across six scenarios, PROP generally improved graduation and futility decisions, reduced inappropriate graduation, and supported the addition of promising regimens compared with dose-based alternatives. It also more accurately estimated regimen-specific toxicity and arm-specific efficacy, while the model-averaging framework favored the efficacy model consistent with the data-generating mechanism. Dose-based approaches performed better for safety stopping in some scenarios, despite less accurate characterization of the regimen--toxicity relationship. PK/PD-informed platform designs can improve adaptive regimen selection and knowledge generation when dose alone cannot adequately characterize treatment regimens.
Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi
Fri 25 Sep
stat.MLcs.LG
pdf html
- interest Language models as statistical objects 0.79
- favorites near “Why Does Deep Learning Work? - A Perspective From Group Theory” 0.78
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.75
Abstract
Deep clustering is a powerful approach for discovering meaningful structures in high-dimensional data by learning a low-dimensional latent representation prior to clustering. Despite its empirical success, assessing the statistical reliability of the resulting clusters remains challenging. Testing discovered clusters on the same data induces selection bias and invalidates classical $p$-values. Selective inference (SI) provides a principled framework for correcting this bias, but existing methods focus on clustering performed directly on the observed features. In this work, we develop an SI framework for deep clustering with a fixed pretrained encoder. The key challenge is that cluster assignments are determined through a nonlinear transformation from the original data space to the latent space, resulting in a substantially more complex selection process than in conventional clustering. Our method provides a computationally tractable way to account for this process and enables valid statistical testing of differences between clusters identified in the latent space. Synthetic experiments demonstrate that the proposed method controls the Type I error rate while achieving higher power than valid but conservative baselines, and genomic applications show that it can identify significant cluster differences while appropriately accounting for selection bias. Our framework provides a principled approach to quantifying the statistical reliability of structures discovered by deep clustering.
Ziqi Qin, Hong Ye Tan, Ander Biguri, Jingwei Liang, Carola-Bibiane Schönlieb
Tue 29 Sep
math.OCcs.NAmath.NAstat.CO
pdf html
- reading operator, processes, statistical interpretation, standard
- interest Kernels, operators and functional data 0.75
- favorites near “Monte Carlo Geometry Processing” 0.76
Abstract
Plug-and-Play (PnP) algorithms are a class of iterative methods for inverse imaging. Within an optimization algorithm, they combine a flexible fidelity term, encoding the forward operator, and a pretrained image denoiser, in order to deal with more severe corruptions such as blurring or downsampling when reconstructing an image. This work studies provably convergent PnP methods for mixed-noise forward processes by using the infimal convolution as a fidelity term, providing a statistical interpretation as a joint maximum a-posteriori estimator over both noises, and preserving the Bayesian MAP interpretation of PnP methods. Independently, we extend the PnP formulation to multiple prior terms using the Davis--Yin three-operator splitting. This extension can be combined with either standard fidelities or the proposed mixed-noise infimal-convolution fidelities. We verify that these generalized PnP methods are convergent under standard Kurdyka--Lojasiewicz conditions. Numerical experiments on Laplace-Gaussian and Poisson--Gaussian noise demonstrate stable convergence of single-prior and multiple-prior PnP methods, and divergence under fidelity mismatch. Furthermore, PnP with infimal convolution fidelities are able to scale to noise up to 38% standard deviation, with multiple priors reaching different stationary points that qualitatively preserve more textural properties.
Yanjin Xiang, Zhihua Zhang
Thu 24 Sep
math.STmath.PRstat.MLstat.TH
pdf html
- reading state, state space, inequality, kernels
- favorites near “Spectral State Space Models” 0.73
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.69
Abstract
Let $P$ be an irreducible reversible Markov kernel on a $m$-state space $\Omega$, and denote its right spectral gap $\gamma=1-\lambda_2(P)$. From a stationary trajectory, let $K_t$ be the occupation-count vector of the length-$n$ window beginning at time $t$. The stationary pair $(K_0,K_1)$ defines a reversible projected count kernel $\widetilde P_n$. For every $m\ge2$, let \[ c_m^\star= \inf_{\substack{ P,\; n \ge 2}} \frac{n\Gap(\widetilde P_n)}{\Gap(P)}. \] We prove \[ \frac1{1080m}\le c_m^\star\le q_{m-2}, \qquad q_0=\frac14,\quad q_{r+1}=q_r(1-q_r). \] We also show $q_{m-2}=(m+\log m+O(1))^{-1}$, which implies that $c_m^\star=\Theta(m^{-1})$. Thus, the optimal comparison coefficient has order $m$, although its exact value remains open. The lower bound also holds for $n=1$ and is uniform in $P$, including sparse and periodic kernels. Its proof combines short-window decorrelation with an averaged anchor-excursion decomposition, a Green-kernel hitting estimate, and a stopped Carleson--Hardy inequality. A nested rare-state construction produces finite $m$-state witnesses whose normalized Rayleigh quotients approach $q_{m-2}$ through an ordered sequence of limits. For every fixed finite irreducible reversible aperiodic kernel on at least two states, $\Gap(\widetilde P_n)=\Theta_P(n^{-1})$.
Jaehoan Kim, Simon Mak
Tue 29 Sep
stat.MEmath.STstat.COstat.MLstat.TH
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.74
- reading log, basis, posterior inference, dimensionality
- favorites near “Recent Developments in Machine Learning Methods for Stochastic Control and Games” 0.75
Abstract
Virtual simulators are widely used for studying complex physical phenomena, from particle collisions to rocket propulsion. Such "computer experiments" can be highly time-intensive, and a Bayesian surrogate model can be used for efficient emulation with reliable uncertainty quantification. To train accurate surrogates with a limited sample size $n$, recent work has explored the incorporation of monotonicity (or isotonicity) information, which can often be elicited from physical systems. In practical applications with many input variables, however, existing Bayesian isotonic models can face statistical and computational limitations, which may result in worse performance compared to models that do not incorporate isotonicity. We propose a new transformed additive isotonic model (TAIM), which aims to tame this "curse-of-dimensionality". TAIM makes use of a flexible transformed additive isotonic modeling framework, which leverages a data-estimated link transformation and a monotone basis model with spike-and-slab priors on basis weights. Prediction-wise, TAIM achieves (up to log factors) a posterior contraction rate of $O(n^{-1/3})$ when the true black-box function is in a transformed additive isotonic form with mild smoothness conditions. Such a rate does not depend on the input dimension $d$ for terms involving $n$, which softens the effect of dimensionality on posterior predictions. Computation-wise, TAIM allows for efficient posterior inference via a carefully designed Gibbs sampler, where each sampling iteration requires only linear work in $d$. We further present an extension of TAIM that can model potential deviations from transformed additivity. Numerical experiments and two applications show the effectiveness of TAIM for isotonic surrogate modeling with many input variables.
Krishnakumar Balasubramanian, Zhaoyang Shi
Wed 30 Sep
cs.LGcs.AIstat.ML
pdf html
- reading manifolds, memory, memories, curvature
- interest Kernels, operators and functional data 0.76
- favorites near “Hypoelliptic diffusion maps I: tangent bundles” 0.80
Abstract
Geometry does more than constrain an associative memory: curvature determines what it remembers and which states it creates. We develop intrinsic dense associative memories on Riemannian manifolds by casting memory as Epanechnikov kernel-density mode seeking. We compare geodesic and volume-corrected energies and show that curvature separates their behavior. We prove that geodesic memory always retains an isolated pattern, while corrected memory obeys a sharp Ricci-curvature threshold: positive curvature can erase memories in high dimensions, while negative curvature reinforces them. We derive geodesic capacity scalings of $q_\beta^{-1/2}$ for retaining every pattern and $q_\beta^{-1}$ for a typical one, where $q_\beta$ is the pairwise kernel-overlap probability. We show how overlap \emph{creates} novel memories: designed $N$-pattern configurations realize all $2^N-1$ subset modes, but random data at the storage threshold yield only a Poisson number. We establish exact one-step recall using Riemannian mean shift. In simulations, we recover the predicted curvature transition and every designed mode. On WordNet's full noun hierarchy, we demonstrate that volume correction improves low-capacity retrieval. Together, our work shows that curvature is a design variable for associative memory, not merely a property of the data.
Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari
Fri 25 Sep
stat.MLcs.LG
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.76
- favorites near “The Consensus Game: Language Model Generation via Equilibrium Search” 0.80
- interest Language models as statistical objects 0.75
Abstract
We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5\%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.
Montserrat Fuentes, Veronica B. Patterson
Thu 24 Sep
stat.MEstat.AP
pdf html
- reading log, functional, complicated, spectrum
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.73
- interest Kernels, operators and functional data 0.75
Abstract
Autism spectrum disorder is a heterogeneous neurodevelopmental condition whose functional brain organization varies across individuals and imaging centers. Resting-state functional connectivity provides an opportunity to study this variation, but analysis is complicated by high dimensionality, strong dependence among connections, and substantial site heterogeneity. We introduce a Hybrid Spatial Statistical Learning Framework for individualized probability estimation in multisite autism neuroimaging. The framework combines edge-level connectivity, graph-based network organization, and participant characteristics as complementary representations of a common probability target. Fine-scale connectivity retains discriminatory information, while broader representations stabilize probability estimates under site shift. All preprocessing and model development are performed within complete site-held-out validation to assess transportability to unseen acquisition environments. Simulations show that the hybrid preserves the discrimination of the strongest edge model while improving probability accuracy as between-site heterogeneity increases. In 860 participants from 20 ABIDE sites, the hybrid retained ROC AUC while reducing Brier score and log loss. These results show that multiscale integration can improve the reliability and transportability of individualized probabilities from complex, spatially dependent biomedical data.
Guozheng Dai, Yiyun He, Ke Wang, Yizhe Zhu
Fri 25 Sep
math.PRmath.STstat.TH
pdf html
- interest Computational methods for heterogeneous agent and dynamic economic models 0.75
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.74
- reading moment, inequalities, random variables, improvement
Abstract
We establish moment and mixed-tail inequalities for fixed-order decoupled homogeneous chaoses generated by independent, centered, sparse bounded random variables. Our bounds apply to arbitrary real rectangular coefficient tensors and describe the fluctuation scales through weighted slice and partition norms, with a Bennett-type logarithmic improvement in the largest-entry term. As an application, we derive guarantees for sparse Khatri--Rao embeddings that explicitly account for sparsity and input geometry.
Peiheng Gao, Chen Yang, Shimin Zhang
Tue 29 Sep
cs.CLcs.LGstat.AP
pdf html
- your work near “Estimating Distributional Responses to Macroeconomic Shocks” 0.71
- interest Macroeconometrics and time series 0.77
- favorites near “Hicks-Arrow Prices for US Federal Debt 1791-1930” 0.71
Abstract
Sentiment analysis is a common tool for converting unstructured text into quantitative signals in finance and risk management. Yet most applications reduce the output to a discrete polarity label or a single predictive feature, overlooking the distributional structure of sentiment intensity in consumer complaint narratives. In this paper we treat negative sentiment in consumer complaints as a bounded continuous variable and study its full distribution rather than a single label. We score each narrative with a transformer classifier, model the scores with Beta distributions, and compare the fitted distributions of meritorious and non-meritorious complaints through the Kullback Leibler divergence and the squared Hellinger distance. The fitted distributions are then linked with dollar amounts and company response outcomes to construct anomaly diagnostics that flag complaints whose textual severity is inconsistent with the recorded relief. We find that the two groups have strongly overlapping distributions, so negative sentiment intensity is not a sharp classifier of outcomes on its own; combined with monetary and categorical attributes, it isolates unusually severe complaints for operational risk monitoring. Treating sentiment analysis as continuous distributional measurement, this study links sentiment extraction, bounded response modeling, and anomaly detection in a unified framework for consumer complaint assessment.
Heon Lee
Tue 29 Sep
math.PRcs.ITmath.ITmath.STstat.TH
pdf html
- reading eigenvalue, root, tends, state
- favorites near “Random Walks on Simplicial Complexes and Harmonics” 0.80
- your work near “Automated Solution of Heterogeneous Agent Models” 0.73
Abstract
We study robust reconstruction for the \(q\)-state Potts broadcast process on a Galton-Watson tree in the growing-alphabet regime \(q\to\infty\), with boundary depth growing sufficiently quickly in relation to \(q\). We show that, in the regime \(d\lambda >1\), where \(d\) is the expected number of offspring and \(\lambda\) is the nontrivial eigenvalue of the Potts transition matrix, the root posterior is asymptotically unchanged by a broad class of noise channels applied independently to the boundary labels. This remains true even when the probability of retaining the true label at an individual boundary vertex tends to zero as \(q\to\infty\). More precisely, under suitable moment and noise assumptions, the noiseless and noisy root posteriors converge to one another in expected total variation, and hence have the same asymptotic Bayes-optimal reconstruction accuracy.
Danna L.Cruz-Reyes, Juan Sosa, Carlos A.Martínez
Wed 30 Sep
stat.MEstat.AP
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.69
- interest Bayesian computation and structural estimation 0.79
- your work near “Computational Methods for Economic Models with Function Valued States” 0.75
Abstract
Crash counts on urban road networks exhibit spatial dependence driven by road connectivity rather than Euclidean proximity alone. We propose a Bayesian negative binomial model that represents spatial variation through a parsimonious combination of spectral components defined directly on road segments. The model combines the normalized random edge neighborhood Gaussian (RENeGe) operator with a continuous spike-and-slab prior, with mixture indicators marginalized out. Eigenvalue-dependent slab variances link spectral regularization to network-supported smoothness, while posterior slab-membership probabilities quantify support for individual components and uncertainty in effective spectral complexity. Segment length is used as an exposure offset, so fitted rates describe crash frequency per unit road length rather than traffic-adjusted risk. In a simulation study with 1,200 datasets and 4,800 Bayesian fits, the proposed model achieved the lowest mean spatial-field recovery error and highest mean test log predictive density under sparse RENeGe truth. Under alternative generating mechanisms, its predictive performance remained close to dense spectral models and exceeded that of a nonspatial model. Applications to Barcelona and central Bogota showed distinct posterior spectral patterns. Exploratory leave-one-segment-out comparisons found similar predictive accuracy for the proposed model and a low-rank edge CAR model, both outperforming a nonspatial negative binomial model, while a spectral BYM2 model achieved the highest exploratory predictive score. These results support probabilistic summaries of spectral complexity without a detectable predictive penalty relative to the corresponding CAR model.
Junwoo Bae, Jin-Hyun Ahn
Tue 29 Sep
cs.AIstat.ML
pdf html
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.84
- reading makes, auxiliary, exchanges, language
- interest Language models as statistical objects 0.74
Abstract
Split Federated Learning (SFL) enables resource-constrained clients to participate in collaborative training, but vanilla SFL exchanges smashed data and gradients at every batch, which incurs significant communication overhead. Recent methods reduce this overhead with an auxiliary network at the client-side cut layer. However, we identify that this approach makes the client optimize a local objective that differs from the end-to-end objective, which fundamentally limits the collaborative training between the client and the server. We propose Compensated Feedback based SFL (CoeF-SFL), a communication-efficient framework that retains the end-to-end objective without any auxiliary network. In CoeF-SFL, the client and the server exchange the smashed data and the gradients once per round and reuse them during local training. Since this reuse makes the gradients stale on the client side, we compensate them with a curvature-based correction in the activation space and develop two variants. CoeF-D approximates the Hessian with a diagonal gradient outer product, while CoeF-J exploits the tractable Jacobian-based Hessian of a surrogate loss that upper-bounds the true loss. We provide the theoretical background of each method, characterizing its compensation. Across vision and language tasks, model capacities, cut layers, and data distributions, CoeF-SFL significantly outperforms auxiliary-network-based methods under the same communication frequency, and the improvement is most substantial on vision tasks. Code is available at https://anonymous.4open.science/r/CoeF-SFL-2686/README.md
Zhongqi Fan, Tianyou Zhang, Fei Chen
Fri 25 Sep
stat.MLcs.LG
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.75
- interest Language models as statistical objects 0.77
- favorites near “The Consensus Game: Language Model Generation via Equilibrium Search” 0.80
Abstract
Selective classification enables a model to abstain from predictions on uncertain instances, but existing approaches typically reject them through confidence scores, predefined coverage constraints or instance-level distance measures. These approaches may overlook the collective geometric structure of difficult samples in learned representation spaces. We propose Guided Clustering-based Uncertain Learning (GCUL), a geometric-guided selective classification framework that identifies misclassified and ambiguous instances as a potential confusion attractor in the representation space. GCUL uses a three-phase procedure to initialize, cluster, and explicitly relabel this uncertain region, allowing the rejection boundary to emerge from the underlying representation geometry rather than from a prescribed rejection rate. We further derive a selectivity score and a geometric sufficient condition that characterizes when rejection can provide positive operational utility, enabling pre-deployment feasibility assessment. GCUL improves DistilBERT accuracy from 89.37 percent to 94.98 percent with less than 9 percent rejection. Beyond accuracy, our selectivity score correctly pre-detects the only dataset (GoEmotion) where all baselines fail, and controlled simulations yield 6.1 percent Type-I and 0 percent Type-II errors, validating the sufficient condition's conservatism. These results suggest that collective representation geometry provides a useful alternative perspective for selective prediction.
Linjian Meng, Siyuan Gan, YuHan Li, Xiran Wang, Ziyang Ding, Ditang Gou, et al.
Tue 29 Sep
cs.CLcs.LGstat.ML
pdf html
- favorites near “Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs” 0.79
- reading rich, language model, policy, large
- interest Language models as statistical objects 0.73
Abstract
On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student via rollouts generated by the student's policy. However, estimating the gradient of the reverse KL divergence in OPD remains a challenge. Using only the sampled token from the student-generated rollout is computationally cheap but provides limited distributional supervision, which will degrade accuracy. In addition, using the full vocabulary provides complete distributional supervision but is computationally expensive. Therefore, recent works propose Top-$k$ OPD (TK-OPD) that use selected top-$k$ tokens, which provides richer distributional supervision than sampled-token estimation at substantially lower computational cost than full-vocabulary estimation. Unfortunately, using only the selected top-$k$ tokens induces bias, leading to accuracy degradation, as the probability mass outside the selected top-$k$ tokens is discarded. To address the bias of TK-OPD, we propose Tail-Corrected Top-$k$ On-Policy Distillation (TT-OPD). It preserves the advantages of TK-OPD, including rich distributional supervision and low computational cost, while providing an unbiased estimator of the gradient of the reverse KL divergence. The key insight of TT-OPD is to use not only the selected top-$k$ tokens, but also the sampled token from the student-generated rollout, thereby recovering the discarded probability mass in expectation, avoiding the bias. Experimental results demonstrate that TT-OPD significantly outperforms other tested OPD variants.
Xiaopeng Xiang, Wenlin Dai, Marc G. Genton, Wenjia Wang
Mon 28 Sep
stat.ME
pdf html
- interest Kernels, operators and functional data 0.77
- reading gaussian process, large, large datasets, make
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.70
Abstract
Gaussian process prediction is a central tool in spatial statistics, but standard implementations require dense matrix operations that become prohibitive for large datasets. We propose a scale-adjusted compactly supported working correlation for Gaussian process prediction, using the generalized Wendland family with a support radius $\phi_n$ that is allowed to decrease with the sample size $n$. Under fixed-domain asymptotics with quasi-uniform designs, a fixed support radius does not yield asymptotic sparsity, whereas a shrinking support radius can make the covariance matrix sparse. We show that $\phi_n$ and the regularization parameter can be jointly chosen so that the resulting predictor preserves the optimal integrated mean squared prediction error rate while reducing the number of nonzero covariance entries and the associated sparse matrix-vector multiplication cost. We also establish an analogous rate-preserving sparsification result for kernel ridge regression. Simulations and an ERA5 temperature application show that the proposed method achieves competitive prediction accuracy while retaining the computational efficiency provided by sparse linear algebra.
Phillip Lo, Sudarshan Babu, Dari Kimanius, Aly A. Khan
Mon 28 Sep
cs.AIstat.ML
pdf html
- reading grids, implicit, bank, large
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.79
- interest Language models as statistical objects 0.70
Abstract
CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural representations (INRs) are able to model volumetric data as scale-agnostic, coordinate-conditioned functions. INRs are therefore attractive for cryoEM, but fitting a separate INR for each map is too expensive for large-scale feature extraction and produces representations that are not aligned across samples. We introduce Atelier, a self-supervised framework that amortizes INR fitting for reconstructed cryoEM maps. Pretrained on 5,439 Electron Microscopy Data Bank maps, Atelier is a transformer-based hypernetwork that generates high-fidelity reconstructions across a wide range of protein structures, including large multi-subunit assemblies. Beyond reconstruction, the INR generated by the pretrained transformer exposes a continuous, local feature field through its intermediate activations at any spatial query point, a property that voxel grid and patch-tokenizer architectures do not naturally provide. Used as auxiliary channels to a 3D nested U-Net annotation head trained from scratch, these coordinate-conditioned features improve performance on eight voxel-level property prediction tasks over a volume-only baseline. Our results demonstrate that amortized implicit neural representations are an effective primitive for geometry-aware analysis of cryoEM data.
Quan D. Bui, Nguyen Do, An Nguyen Dang, Huyen Nguyen, Nhu Duc Minh Nguyen, My T. Thai
Tue 29 Sep
stat.MLcs.AIcs.LG
pdf html
- interest Language models as statistical objects 0.79
- reading bases, structural, basis, sacrificing
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.76
Abstract
Existing interpretable graph additive models still face limitations in either computational scalability or modeling flexibility. In terms of structural modeling, previous approaches either face quadratic scaling costs or sacrifice explicit source-to-target contribution decomposition. In terms of feature components, they rely either on per-feature neural networks or on single shared bases with limited feature specialization. We address both problems by introducing HARMONIA: Interpretable Graph Learning through Mixtures of Neural Bases, an interpretable-by-design framework. For feature modeling, HARMONIA introduces a Mixture of Neural Bases (MoNB), which routes features to specialized basis experts, enabling parameter sharing without sacrificing feature-specific specialization. For structural modeling, HARMONIA uses Relative Random Walk Probabilities (RRWP) to capture multi-hop and multi-path relationships, and proposes Sparse RRWP Aggregation (SRA) to compute these interactions through sparse graph propagation without quadratic pairwise complexity. HARMONIA retains a simple additive form in which predictions decompose into feature responses modulated by structural influence. Empirically, HARMONIA achieves stronger explanation recovery than existing interpretable graph baselines while maintaining competitive predictive performance and scaling to graphs with millions of nodes. These results show that interpretable graph learning can remain both faithful and scalable without sacrificing predictive effectiveness.
Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai
Mon 28 Sep
cs.LGstat.AP
pdf html
- your work near “Timing as an Action: Learning When to Observe and Act” 0.76
- reading units, gains, explains, explain
- favorites near “Stabilized Neural Prediction of Potential Outcomes in Continuous Time” 0.76
Abstract
Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration procedure for mortality prediction under cross-hospital distribution shift. It selects a measurement on an independent sample, groups patients by whether that measurement is recorded, and applies Mondrian calibration within each group, so no calibration outcome is reused. We evaluate the procedure across hospitals in eICU and across care units within one MIMIC-IV hospital, using three predictors. Relative to pooled calibration, it reduces the average worst-group coverage gap on its selected groups in all six settings, with a median reduction of 1.9 percentage points; paired site-bootstrap intervals exclude zero in five. These gains do not extend uniformly. Calibration by predicted risk achieves smaller gaps on a broader panel of missingness groups, and when eICU hospitals are evaluated separately, the gain shrinks for all three predictors and reverses in sign for one. We explain this discrepancy with a hospital-level decomposition. Pooling reweights hospitals through a covariance between group shares and coverage errors, and lets errors of opposite sign cancel: weighting explains the reversal, and cancellation accounts for most of the attenuation for the other two predictors. Constructed population distributions show that pooled and within-hospital evaluations can rank calibration methods oppositely even without sampling noise. Pooled improvement alone therefore cannot establish better coverage within hospitals, even when the calibration groups are fixed.
Siriac Seboka, Laurent Briollais, Yun-Hee Choi, on behalf of BCFR, Lajmi Lakhal-Chaieb
Mon 28 Sep
stat.ME
pdf html
- reading history, women, cancer, personal
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.74
- favorites near “Graphical criteria for the identification of marginal causal effects in continuous-time survival and event-history analyses” 0.80
Abstract
Breast cancer risk assessment and prediction models are very important tools for the clinical management of healthy women carrying pathogenic variants in known cancer genes. We propose in this paper a comprehensive prediction model for the personalized clinical management of women with pathogenic variants in {\it BRCA1}. This model is able to estimate the risk of BC accounting for the full history of this cancer within the family, the personal and family history of risk-reducing salpingo-oophorectomy as well as the exact age of this intervention among family members. The effect of oophorectomy on the risk of developing breast cancer is evaluated by estimating and conducting inference about the regression coefficients in a Cox model with time-varying covariates. We model the within-family dependence in the ages at onset of cancer using a Gaussian copula whose correlation matrix accommodates the different pairwise family relationships. We develop an iterative algorithm to estimate the parameters of the considered model in the presence of a selection bias. We evaluated the performance of the proposed cancer risk estimation method from family data by simulations and illustrated its use through an application to the breast cancer family registry.
Mitushi Goyal, Tarun S., Riddhanya Senapathi, Arun Raman
Wed 30 Sep
cs.LGcs.DCstat.ML
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.76
- interest Kernels, operators and functional data 0.72
- favorites near “Data Science at the Singularity” 0.78
Abstract
Federated clustering methods that do not require the global number of clusters $K$ still assume that each client knows its local number $K_g$. This assumption is hard to justify when clients know no more about their data than the server does, as in fault diagnosis across independently operated industrial sites. We propose a two-phase framework in which neither count is known: each client first estimates $K_g$ from its own data, and an aggregator that requires local counts, such as FedGEM, then uses these estimates in place of the true values. For the first phase we introduce Adaptive Split--Merge (ASM), which grows a spherical Gaussian mixture by BIC-driven splitting and then merges excess components. ASM uses no labels, selects its hyperparameters on held-out client data only, and makes no assumption about how clusters are shared across clients. We derive a closed-form split criterion whose critical cluster size falls with anisotropy and rises with dimension, and show empirically that over-fragmentation grows with the number of points per cluster, which federation divides among clients. Across eight datasets, ASM with FedGEM attains a mean ARI of 0.333, against 0.256 for the next best label-free estimator and 0.361 when the true local counts are supplied. It also gives the most reliable global estimates of $K$ and is robust when client size is decoupled from local cardinality.
Xinyu Song
Fri 25 Sep
math.STstat.TH
pdf html
- reading minimax risk, argument, state, moment
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.71
- interest Adaptive data collection and anytime-valid inference 0.76
Abstract
We establish a lower bound for estimating the von Neumann entropy from independent outcomes of any fixed rank-one POVM. A rotation-averaged van Trees argument gives a global minimax risk of at least $(d/n)\log^2\{n/(4d)\}$ when $d\ge C$ and $n\ge Cd$, without a projective-design assumption. We also characterize risk on an operator-norm ball of radius $r$ around the maximally mixed state. We allow an approximate second moment: on the trace-zero Hermitian subspace, the measurement frame may differ by $\varepsilon<1$ from the tight projective frame. A clipped estimator based on canonical dual shadows and the complete U-statistic for purity has risk at most $d^3r^2/n+d^4/n^2+d^6r^6$. Lower bounds under the same frame control yield the local minimax rate $d^3r^2/n+d^4/n^2$ when $n\ge Cd^2$ and $r$ lies in an explicit matching range. For every fixed upper bound on $\varepsilon$, approximation changes only the constants, not the powers of $d,n,r$. At the critical radius $r=n^{-1/2}$, the local risk is asymptotically negligible relative to the global risk when $n\log^2(n/d)\gg d^3$. This separation holds for projective 2-designs, including global Clifford measurements in qubit dimensions, and for their uniformly well-conditioned frame approximations. Finite-sample experiments in dimension four illustrate the critical-radius benchmark and the effect of a nonexact frame.
Anton Sokolov, Paolo Giudici, Vasily Kolesnikov
Fri 25 Sep
stat.ME
pdf html
- interest Adaptive data collection and anytime-valid inference 0.80
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.73
- favorites near “Data Science at the Singularity” 0.78
Abstract
Artificial Intelligence trustworthiness scores are moving from research dashboards into compliance claims and procurement decisions. A compliance claim is a statistical decision based on the comparison between an estimated compliance metric and a set threshold. The recently proposed integrated SAFE AI metrics are based on three basic rank-graduation measures of accuracy, explainability and robustness, expressed on a common footing. The metrics are integrated into a single compliance score, so far reported as a point value. In this paper we propose to add an uncertainty layer to the SAFE AI compliance score. All component metrics are estimated on a shared test sample, so their errors covary. We estimate the full covariance of the component vector with a paired bootstrap, and quantify departures from the source-wise independence baseline suggested by the integrated-metrics framework's variability decomposition. A closed-form identity splits such departure into an uncertainty-weighted effective dimension and an exposure-weighted data-driven correlation, with both depending on the aggregator and the estimated uncertainty of the component metrics. The application of our proposal to both real and simulated data shows that the assumption of independence can either narrow or widen confidence bands, while the direction and magnitude of the effect depend on the aggregator type, machine learning model, and perturbation family. The resulting confidence region supports compliance decisions, which can be documented in a formal uncertainty certificate.
Han Chen, Yingrui Li
Wed 30 Sep
cs.LGstat.ML
pdf html
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.76
- interest Bayesian computation and structural estimation 0.78
- favorites near “Inference with Hamiltonian Sequential Monte Carlo Simulators” 0.71
Abstract
A probability used for a decision should refer to the same event across equivalent requests. We introduce probability contracts, a benchmark connecting exact finite-world posteriors, validated event transformations, and failure-aware decision evaluation. Four model-interface configurations are evaluated on 1,000 worlds. Their assessments differ across accuracy, coherence, and decision loss: Kev has lower aggregate canonical posterior error than Jev, but larger complement and coarsening residuals, with accuracy ordering varying by stratum. Jev's Event and Choice interfaces induce different binary actions on 32.8% of valid pairs at defer cost 0.10. A post-hoc analysis finds that disagreement certifies only 11-52% of mean binary pair error across configurations. An elementary action-region characterization explains when averaging changes decision loss relative to randomly selecting one interface. Although averaging cannot worsen that baseline's expected Brier score, its decision effect depends on the cost and crossed boundaries; the observed same-baseline penalties occur in configurations already worse than always deferring. The benchmark makes these distinctions measurable without treating consistency as accuracy or a score improvement as a decision guarantee.
Thomas B. Michelon, Avidane C. Caballero, Hsieh Fushing
Thu 24 Sep
stat.COq-bio.QM
pdf html
- reading sex, data analysis, interacting, heterogeneity
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.70
- interest Empirical Bayes, shrinkage and decision theory 0.73
Abstract
We employ Computational Taxonomy (CT) to reanalyze the penguin data set penguins_lter by validating and addressing two biological issues: Sexual Size Dimorphism (SSD) and mate-selection criteria. Via Scientific Data Analysis (SDA) computing, CT constructs a Taxonomic Hierarchy by splitting Species first and then Sex, without involving Island, to achieve less complexity. This Taxonomic Hierarchy validates SSD as a branch comparison: (Species, Sex = Male)-vs-(Species, Sex = Female), upon which SDA explores all potential pieces of associative information from all covariate feature-sets, including interacting effects from order-2 to order-4, and then confirms them via their idiosyncratic reliability checks. The collective of confirmed information pieces are displayed on a heatmap platform to manifest underlying dynamics of SSD with explicit block-structured heterogeneity found within males and females. SSD dynamics is explained through mechanistic dependence pertaining to one chief factor consisting of up to 8 feature-sets: Body-Mass coupled by combinations of {Culmen-length,Culmen-depth, Flipper-length}, and two minor factors consisting of low-order combinations of {Culmen-length,Culmen-depth, Flipper-length}. Such Intra-Sex heterogeneity invalidates all Logistic regression modeling on SSD in the original paper. Further, we explore potential mate-selection criteria through the data-frame of Nest-ID within-species homogeneity.
Yuren Hao
Wed 30 Sep
cs.LGcs.ETstat.ML
pdf html
- reading equilibrium, analog, control, computing
- interest Computational methods for heterogeneous agent and dynamic economic models 0.71
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.65
Abstract
Physical neural networks and analog in-memory computing could reduce the energy cost of neural network training. Realizing this potential, however, requires optimizers that combine effective learning with physical implementability. SGD fits local analog updates but struggles on transformers, while Adam family is unstable against analog bias. Muon offers strong training performance, but its Newton--Schulz orthogonalization relies on dense matrix-matrix products. To address this obstacle, we introduce Physical Muon, which computes the orthogonalization as the equilibrium of a continuous-time flow. Random probes approximate the flow using matrix-vector products, reciprocal reads, and local rank-1 writes. To test whether this replacement preserves training performance, we evaluate it on a 10.95M-parameter transformer. The dense flow's mean validation cross-entropy is 0.0085 above Newton--Schulz across nine seeds per method; the probe implementation is 0.0188 above the control across two seeds. Circuit simulations further reproduce the flow dynamics and yield comparable training behavior.
J. Derek Tucker, Gavin Collins, Gabriel Huerta, Justin L. Brown
Thu 24 Sep
stat.ME
pdf html
- favorites near “Inference with Hamiltonian Sequential Monte Carlo Simulators” 0.76
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.70
- interest Bayesian computation and structural estimation 0.76
Abstract
Bayesian calibration of functional-output computer models typically relies on dimension reduction techniques, such as functional principal component analysis, which assume that differences among simulator realizations arise only from amplitude variation. When simulator output also exhibits phase variation such as shifts in the timing or location of key features, this assumption is violated. Recent work has addressed this issue through elastic calibration, which aligns functional computer model realizations with observed experimental data prior to dimension reduction. Separately, multi-fidelity methods reduce the cost of calibration by supplementing a small number of expensive high-fidelity simulator runs with a larger ensemble of cheap low-fidelity runs. This is typically done through either a mapping strategy, which corrects low-fidelity predictions toward high-fidelity output, or a fusion strategy, which builds a shared basis across both fidelities. This paper combines these two lines of work, introducing elastic multi-fidelity Bayesian model calibration, which aligns high- and low-fidelity functional output to a common reference before applying multi-fidelity mapping or fusion. On a synthetic two-dimensional calibration problem and a dynamic material properties equation-of-state problem, both elastic multi-fidelity strategies match or improve on the leave-one-out predictive accuracy of a mono-fidelity elastic emulator, with the fusion approach achieving the lowest error. Both strategies also produce tighter calibrated posteriors than the mono-fidelity baseline, with the fusion approach providing the best coverage and parameter estimates closest to the true values.
Ali R. Hashemi, Mohammad R. Hashemi, Pavel B. Ryzhakov
Wed 30 Sep
physics.flu-dynphysics.comp-phphysics.data-anstat.ML
pdf html
- interest Kernels, operators and functional data 0.77
- favorites near “Monte Carlo Geometry Processing” 0.77
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.69
Abstract
A droplet breaks up in tens of microseconds, and a recording captures perhaps a dozen frames. The states in between cannot be recovered without repeating the experiment, and simulating them is too costly to sweep an operating envelope. Yet they are present in the corpus as a whole: a campaign spanning a device's actuation range produces morphologies resembling those any single recording missed. Exploiting that requires a representation that is low-dimensional, invertible, and faithful to shape rather than to sampling. Proper orthogonal decomposition supplies the first two but measures distance in sampled coordinates; manifold learning supplies the third but no map back to a shape; elastic shape analysis supplies a shape metric but no reduced coordinates. SHROM composes all three. Interfaces are represented by their square-root velocity functions, a neighbour graph is built over that shape space, and an autoencoder is trained to reconstruct while penalising latents in which graph neighbours are not latent neighbours. The demonstration uses 301,539 inkjet droplet contours and four filmed break-up sequences. The graph term does not improve reconstruction. It determines whether position in the latent carries meaning: clustering the latent of an otherwise identical model recovers the shape-space partition at chance level (ARI = 0.063 +/- 0.059), and at 0.781 +/- 0.049 with the term active. Waveform parameters predict the full contour at R^2 = 0.878 +/- 0.026. Negative results are reported in the same terms. The graph metric proved immaterial across five choices, and interpolation error saturates at the reconstruction limit, so a plain autoencoder leads that task. The main limitation is interpolation across a topology change: no component of the regularised latent holds both a single-component and a post-break-up shape, whereas an unregularised autoencoder mixes them freely.
Ngoc-Hai Nguyen, Thuan Nguyen, Prakash Ishwar, Shuchin Aeron
Tue 29 Sep
stat.MLcs.LG
pdf html
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.82
- interest Kernels, operators and functional data 0.73
- reading inverse, choice, pairs, computing
Abstract
Inverse Optimal Transport (OT) based methods for representation learning learn representations such that the global OT coupling between a pair of data marginals in the representation space, concentrates on the positive pairs. This is in contrast to previous methods that primarily focused on pairwise matching. However, these methods $\textit{DO NOT}$ utilize negative pairs and hence are not truly contrastive in their approach. We show that this leads to issues of dimensional collapse and hence degraded downstream performance. To alleviate this, we develop a novel multi-marginal (MM) inverse OT (IOT) contrastive learning (CL) approach called Neg-MMIOT-CL, which learns representations such that the global multi-marginal OT (MMOT) coupling between a triple of data marginals, with respect to a carefully designed ground-cost between triplets of data points in the representation space, concentrates on the anchor-positive-negative $\textit{triplets}$. For a latent class model, we empirically show that Neg-MMIOT-CL alleviates dimensional collapse. Furthermore, for a specific choice of ground cost for all triplets in representation space, we prove that the optimal representation configuration for Neg-MMIOT-CL exhibits equiangular property for within-class and across-class representations, which translates to Neural-Collapse when the representation dimension is larger than the number of classes minus one -- a result that is $\textit{previously established only}$ for pairwise contrastive learning methods. Finally, we propose Neg-IOT-CL-PushPull, that is a computationally efficient alternative to Neg-MMIOT-CL, alleviating the high cost of computing MMOT plans needed during implementation. We apply these methods on both synthetic and real-world datasets and show significant improvements over existing OT-based contrastive learning methods.
Rajarshi Guhaniyogi, Pritam Dey, Krishnendu Chandra, Aaron Scheffler, Bani K. Mallick
Thu 24 Sep
stat.MEstat.APstat.ML
pdf html
- reading functional, functional data, gp, nonparametric
- interest Kernels, operators and functional data 0.74
- your work near “Computational Methods for Economic Models with Function Valued States” 0.69
Abstract
This article is motivated by an imaging application from the Adolescent Brain Cognitive Development (ABCD) study, aiming to predict task-based brain activation maps from t-fMRI using cortical metrics from structural MRI (s-MRI) and brain connectivity data from resting-state fMRI (rs-fMRI). Hierarchical Bayesian modeling is well-suited for integrating diverse imaging data and quantifying prediction uncertainty. However, progress in this field is limited due to challenges in designing joint priors that capture the structures and interconnections between different imaging modalities, along with computational complexity and lack of theoretical assurances. To address these challenges, the article introduces a novel regression framework that treats t-fMRI and s-MRI images as functional data, incorporating additive non-linear effects of both network and functional predictors on the functional response. Specifically, we employ Gaussian process (GP) priors on coefficients related to the functional predictors to capture their intricate functional dependencies with the response. Furthermore, a GP prior is assigned to encapsulate the non-linear nodal effects of the network predictor on the response function. The method is supported by theoretical results on predictive accuracy for the functional response, and is empirically validated through simulation studies and analysis of multi-modal neuroimaging data from the ABCD study.
David Kepplinger, Anand N. Vidyashankar
Tue 29 Sep
stat.MLcs.ITcs.LGmath.ITmath.PRmath.STstat.TH
pdf html
- interest Language models as statistical objects 0.77
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.78
- your work near “Valid Inference with Imperfect Synthetic Data” 0.73
Abstract
Domain Adaptation (DA) often lacks worst-case guarantees, while Distributionally Robust Optimization (DRO) based only on source data centers its ambiguity set at the source law and ignores available target structure. To bridge this gap, we introduce a doubly-anchored DRO framework whose ambiguity set is the intersection of $\phi$-divergence balls centered at the source law and a source-completed target reference law, the latter pairing the target covariate law with the source conditional law. We derive dual-induced adversarial bridge geometries for symmetric and asymmetric divergence pairings, notably introducing a Kullback--Leibler/squared-Hellinger (KL/HD) bridge. This asymmetric formulation yields a Lambert-$W$ geometry in which source-side exponential risk tilting and target-side Hellinger stabilization enter through distinct terms, attenuating, but not bounding, the effect of large likelihood ratios. Furthermore, without imposing covariate shift, we establish finite-sample generalization bounds for the minimizer of a structural, loss-agnostic density-bridge risk under bounded-overlap conditions; these bounds do not apply directly to the loss-aware DRO min--max estimator. We translate our framework into a bridge-weighted Nadaraya--Watson estimator, proving uniform consistency for the source regression function and pointwise asymptotic normality, with target recovery when the source and target regression functions coincide, as under covariate shift. Finally, an empirical evaluation on a domain-shifted Fashion-MNIST dataset illustrates the finite-sample stability of the asymmetric KL/HD bridge under severe synthetic target-covariate corruption.
John Boland, Matthias Fresacher, David Hill, Shijia Jin, Ya Li, Graham Mills, et al.
Wed 30 Sep
stat.MEstat.AP
pdf html
- reading economic, large, series, sector
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.72
- interest Macroeconometrics and time series 0.74
Abstract
The article identifies the desired properties of the reliability risk metric for the Australian Energy Sector. It proposes a comprehensive set of metrics designed to measure various aspects of the reliability risk when a large portion of power generation is composed of variable renewable energy (VRE). The suggested methods aim to balance the tradeoff between economic benefit and risk management, effectively address the tail risk, and interpret the severity of the outage in a direct way. Properties of the considered reliability metrics are investigated by using simulation studies that use time series of expected unserved energy values. Sample distributions, dependency structures and other statistical properties of the metrics are studied. The results suggest new tools and approaches that can be used to increase the efficiency of the Australian Energy Sector.
Yusuf Öztürk, Enes Göktekin, Bengisu Atlı, Akın Öztürk, Zhixiang Wang, Ulas Bagci
Wed 30 Sep
cs.LGstat.ML
pdf html
- reading versus, heterogeneity, closely, avoid
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.73
- favorites near “Data Science at the Singularity” 0.78
Abstract
Industrial predictive maintenance increasingly depends on learning from equipment spread across sites whose sensor data cannot easily be pooled. Federated averaging (FedAvg) solves this with a central aggregation server; gossip learning removes the server, but its behaviour for recurrent failure-detection models has not been measured under controlled conditions. We compare synchronous ring gossip with FedAvg, isolated local training and a centralized reference for a stacked LSTM that detects imminent failure on the NASA C-MAPSS turbofan benchmark. All methods share one open implementation, architecture, initialization, optimizer, data split and training budget, and the primary endpoint uses one terminal window per test engine to avoid the statistical dependence of overlapping windows. On FD001 (five seeds), gossip reached a terminal-window F1 of 89.6 +/- 1.3%, compared with 89.9 +/- 1.1% for FedAvg, 83.6 +/- 6.7% for local training and 93.5 +/- 2.1% for centralized training, while transmitting the same payload as FedAvg without a coordinator. Node models agreed closely but not exactly (1.8% pairwise decision disagreement versus 5.6% without communication). Across FD002-FD004, peer communication improved terminal-window F1 over local training by 13-28 points; gossip matched FedAvg on FD003 and FD004 but was 4.3 points lower on the multi-condition FD002 subset. Simulated message loss, node failure and server outage changed neither method appreciably, whereas larger rings degraded gossip faster. Ring gossip is therefore a practical serverless alternative when data heterogeneity is moderate, and faster-mixing topologies become important as heterogeneity grows.
Yongda Wei, Chen Zhang, Yifei Wang, Xinyu Wang, Bosen Shao, Hanxi Li, et al.
Tue 29 Sep
cs.CVstat.AP
pdf html
- favorites near “Monte Carlo Geometry Processing” 0.79
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.70
- interest Semiparametric inference and debiased machine learning 0.77
Abstract
Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggregate accuracy. We investigate how this image correspondence supports two needs in robustness evaluation: interpreting paired evaluation results and prioritizing samples for physical testing. To interpret paired evaluation results, we fix both sets of prediction records and vary their correspondence within each class. We prove that classwise correct-correct counts give the same sharp bounds on lost acceptance and mean true-class probability decrease among retained-correct inputs as any feasible five-state refinement. Distinguishing persistent from changed wrong answers can further constrain accepted-error transitions, while shared correspondence can establish policy orderings left unresolved by separate cost intervals. To prioritize samples for physical testing, we retain each image's synthetic responses and rank clean-correct images by their mean true-class probability under corruption. Across 44 classifiers, testing the highest-risk 20% finds 67% and 45% of failures under mild screen and print recaptures, versus 58% and 36% for clean confidence and 60% and 37% for an equal-size natural-transformation average. With both probability averaging and an A3Rank scoring adaptation, the tested corruption set yields higher mean failure recall than the natural-transform set; differences between scores depend on the source and budget. Together, these findings show that the value of correspondence depends on the evaluation objective: classwise counts suffice for specified reliability bounds, while image-specific synthetic responses improve the allocation of physical tests within the evaluated pool.
Rong Pan, Yili Hong, Min Xie
Tue 29 Sep
cs.AIstat.AP
pdf html
- your work near “Automated Solution of Heterogeneous Agent Models” 0.75
- favorites near “Learning Scalable Deep Kernels with Recurrent Structure” 0.78
- interest Language models as statistical objects 0.75
Abstract
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning process that produces the output. Average benchmark accuracy measures capability; it does not quantify this broader reliability claim. This paper adapts established reliability engineering methods, from failure definitions and operational envelopes to FMEA, accelerated testing, field monitoring, and reliability growth, to AI systems. A four-level diagnostic framework classifies failures as component, operational-loop, agentic-conduct, or network and governance failures. Test, evaluation, verification, and validation (TEVV), sequential monitoring, and FRACAS create and refresh evidence. SMART provides statistical guidance for measurement, analysis, assessment, and test planning; the NIST AI Risk Management Framework provides organizational guidance for governance, evaluation, monitoring, and mitigation. Three cases illustrate the program: adversarial testing of a convolutional neural network, perception-error propagation, and autonomous-vehicle disengagements. Established reliability engineering provides a usable foundation; new measurements and safety guardrails are still needed as these systems are self-evolving.
Simon Carter, Zeming Kuang, Lilianne R. Mujica-Parodi, Helmut H. Strey
Mon 28 Sep
stat.MLcs.LG
pdf html
- reading researchers, functional, higher, spectrum
- your work near “Solution of Rational Expectations Models with Function Valued States” 0.69
- interest Kernels, operators and functional data 0.72
Abstract
Optimizing fMRI scan duration and spatial resolution is critical for experimental design, yet traditional correlation-based approaches cannot quantify uncertainty or disentangle scanner measurement noise from true neural variability across subjects. Without principled uncertainty bounds, researchers cannot know whether a protocol is long enough to reliably estimate connectivity, or whether between-subject differences reflect biological variation or noise. We present a Bayesian framework modeling BOLD dynamics as coupled Ornstein-Uhlenbeck processes, using Sequential Neural Posterior Estimation to obtain connectivity posteriors while accounting for frequency-independent measurement noise across the BOLD spectrum. Applied to N = 28 healthy controls (55 scans) at 7T using a functional network atlas (65 DMN regions), the framework quantifies uncertainty across its sources: scanner noise, subject variability, and acquisition length. Spatial analysis identifies a mean of 46 voxels per ROI, roughly half of typical region sizes, as sufficient to achieve 90% of asymptotic precision. At the single-subject level, 7T reaches its within-session precision plateau in approximately 7 minutes versus 10 minutes for 3T, a 40% reduction in required scan time, providing the first direct, model-based quantification of the scan-time advantage conferred by higher field strength. At the population level, 3T requires roughly 37 times more per-subject scan time than 7T for the pooled curves to converge, confirming a consistent advantage of higher field strength at every timescale. Together these findings provide concrete, scanner-specific guidance for protocol optimization, with direct implications for reducing acquisition costs and improving the reliability of connectivity-based clinical biomarkers. We provide code enabling researchers to derive these bounds from their own data.
Pamphile T. Roy
Mon 28 Sep
stat.ME
pdf html
- reading net, far, functions, fewer
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.69
- favorites near “Random Walks on Simplicial Complexes and Harmonics” 0.75
Abstract
Quasi-Monte Carlo methods allow computer experiments to be run with far fewer simulations than crude Monte Carlo. A digital net in base two, such as Sobol', is however only balanced when the number of samples is a power of two, and Latin Hypercube Sampling (LHS) accepts any sample size but only controls the one-dimensional margins. This work proposes a space- filling design defined for any sample size, referred to as LAT for Latinized aperiodic tiling. The unit hypercube is cut recursively across its longest edge following the golden section into N cells of equal volume. One point is then placed in each cell, and the margins are made Latin while every point stays inside its own cell. The construction only uses integer splits and costs O(dNlog N). The recursion is shown to follow the Fibonacci word and its aperiodicity is analysed. LAT is assessed with four L2-discrepancies and with the integration error on analytical functions and engineering emulators, and compared to Monte Carlo, LHS, Halton, Sobol' and a rank-1 lattice. LAT is better than Monte Carlo and LHS as soon as the integrand has interactions. Sobol' remains more accurate at the powers of two, but its error is one to two orders of magnitude larger at other sample sizes. The accuracy of LAT does not depend on the sample size. Finally, the cells form a partition of the hypercube for any N. This allows one to refine the design locally, to search for an optimum by splitting cells and to sample non-rectangular regions.
Parastoo Farajpoor, Mohammadreza Narimani
Wed 30 Sep
physics.geo-pheess.IVstat.AP
pdf html
- reading persistence, absolute, greater, closely
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.67
- your work near “Computational Methods for Economic Models with Function Valued States” 0.59
Abstract
Accurate prediction of post-fire snow conditions does not necessarily establish how wildfire changed those conditions. We developed an explainable geospatial artificial intelligence framework combining multisource Earth observations, meteorological information, and matched before-after comparisons to examine seasonal snow persistence following the 2020 Creek Fire in California's Sierra Nevada. Harmonized Landsat Sentinel-2 observations were used to estimate the fraction of clear-sky observations containing snow during October-July for eleven water years, 2016-2026, on a 500 m analysis grid. We matched 3,778 snow-zone cells inside the fire perimeter to comparable unburned controls using terrain and pre-fire snow conditions; absolute standardized mean differences after matching were no greater than 0.047. HLS persistence agreed closely with MODIS, with a mean annual spatial correlation of 0.947. The landscape-average before-after control-impact contrast was +0.0017, with a year-level 95% confidence interval of -0.023 to +0.027. A stronger response emerged in the highest burn-severity class, where observed persistence increased by 0.026 relative to matched controls, equivalent to 2.6 percentage points. Under 5 km spatial cross-validation, XGBoost predicted raw post-fire persistence with an out-of-fold coefficient of determination of 0.811, whereas predictive skill for the fire-adjusted anomaly reached 0.046. Elevation and temperature together accounted for 65.7% of mean absolute model attribution. These findings distinguish a severity-associated optical snow response from the terrain-climate relationships that dominate predictive skill. Combining matched comparisons with explainable GeoAI provides a practical framework for forest monitoring that separates accurate environmental mapping from inference about disturbance effects.
Negar Soleimani, Pierre Barbillon, Kaniav Kamary, Merlin Keller
Thu 24 Sep
stat.ME
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.75
- interest Bayesian computation and structural estimation 0.75
- favorites near “Geometric MCMC for Infinite-Dimensional Inverse Problems” 0.78
Abstract
When computer codes model complex physical systems, calibration alone is insufficient; one must also assess whether a discrepancy term is needed. In this paper, we study computer code validation through a Bayesian mixture approach that compares a pure-code model with a discrepancy-corrected model. The method relies on the posterior distribution of a mixture weight, which measures the relative support of the two competing distributions. Under the assumption that the code is linear in the calibration parameters, or can be well approximated by a linear surrogate, we show that mixture component-shared parameters can be used to combine flexible modeling, even when noninformative priors are assigned to some common parameters. Inference is performed using a Metropolis-within-Gibbs algorithm. In addition, we introduce a thresholded allocation rule that complements the global mixture weight by providing a local diagnostic of where the discrepancy-corrected component is truly needed along the input domain. Beyond global model comparison, the proposed approach is also able to perform local model discrimination by identifying where the discrepancy-corrected component is truly needed along the input domain.
Arlina Shen, Dehua Bi, Ruben P.A. van Eijk, Lu Tian, Ying Lu
Thu 24 Sep
stat.ME
pdf html
- reading functional, treatment, conceptual, questions
- your work near “Estimating Treatment Effects with Observed Confounders and Mediators” 0.74
- interest Causal inference and identification in econometrics 0.72
Abstract
Hierarchical composite endpoints combining survival and longitudinal functional outcomes are increasingly used in clinical trials, especially when death precludes subsequent functional assessment. The Finkelstein--Schoenfeld strategy analyzes such endpoints through prioritized pairwise comparisons, with survival compared before function. In amyotrophic lateral sclerosis (ALS), this strategy is implemented in the Combined Assessment of Function and Survival (CAFS), which combines survival with the ALS Functional Rating Scale Revised. Although such endpoints provide a clinically meaningful summary of overall treatment benefit, investigators may also want to understand whether the treatment effect is driven by survival, functional outcome, or both. Motivated by the estimand framework, we use ALS-informed simulations from a joint longitudinal--survival model to study settings in which treatment effects on survival and function align or conflict. The simulations show that decomposing the composite win probability into survival and survivor-based functional contributions clarifies their relative roles and separates the composite treatment-benefit question from function-focused questions, including the while-alive comparison and conceptual alternatives based on hypothetical and always-survivor estimands. The functional win probability underlying the while-alive comparison can be subject to survivor-selection bias when treatment affects survival, and inverse-probability weighting can attenuate selection induced by measured predictors under appropriate assumptions. Brief supporting analyses based on principal stratification and multiply robust estimation illustrate the always-survivor estimand. This framework provides practical guidance for reporting and interpreting hierarchical composite endpoints, with CAFS in ALS serving as a concrete motivating example.
Tshemollo Rapolai, Seite Makgai, Mohammad Arashi
Fri 25 Sep
stat.MLcs.LG
pdf html
- reading causal, feature spaces, double, heterogeneity
- favorites near “Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression” 0.79
- interest Language models as statistical objects 0.70
Abstract
Predicting blood-brain barrier (BBB) permeability is critical for central nervous system drug discovery. Using the MoleculeNet BBBP dataset (n = 2039), this study systematically ablates molecular feature spaces to isolate featurisation from model architecture. We evaluate three feature families (Morgan fingerprints, RDKit physicochemical descriptors, SMILES bigrams) across four learning algorithms. Results demonstrate that predictive performance depends jointly on feature representation and algorithm. Dynamic Random Forest using combined features achieved the highest mean AUC (0.970, 95% CI: 0.963-0.977). Second, this optimal representation enables exploratory estimation of heterogeneous associations between molecular structure and BBB permeability using Generalized Random Forests. Constructing a pseudo-treatment from a LogP median split, we applied double/debiased machine learning to account for confounding. Orthogonalization substantially attenuates the heterogeneity detected by naive causal forests; no conditional effects remained significant after false discovery rate correction (smallest adjusted p = 0.082). Furthermore, orthogonalized feature importance shifted toward residual structural information in SMILES bigrams. Ultimately, once observed confounding is properly accounted for, evidence that LogP-BBB associations vary systematically across chemical space is insufficient. This underscores that feature representation and model architecture are coupled design choices, and that unorthogonalized causal forests risk overstating genuine treatment effect heterogeneity.
Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees, Chiara Vanalli, et al.
Mon 28 Sep
cs.LGq-bio.PEstat.ML
pdf html
- favorites near “The Geography of Development: Evaluating Migration Restrictions and Coastal Flooding” 0.69
- reading making, know, consistent, hundreds
- your work near “Valid Inference with Imperfect Synthetic Data” 0.69
Abstract
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds of millions of community-science records. At this scale, averaging performance hides substantial species-level variability, particularly for rare species, often of greatest conservation concern. Records are also strongly biased, making occurrence counts misleading. Accounting for these factors is essential for a reliable and informative evaluation of multi-species SDMs. Here, we introduce a Sampling-Aware Global Evaluation (SAGE) benchmark, combining GBIF records for training with sPlotOpen vegetation plots for presence-absence evaluation across 5771 plant species. We propose an evaluation framework that groups species based on two properties, sampling effort and relative prevalence, which describe how densely a species' range is sampled and how frequently the species is recorded. Evaluating single-species SDMs and multi-species DeepSDMs, we find that Random Forests and DeepSDMs perform best overall, but neither dominates: DeepSDMs outperform single-species SDMs for infrequently recorded species while offering no consistent advantage for well-sampled ones. Crucially, this advantage emerges only when established bias-correction practices, such as spatial thinning and reweighting, are carried over to the deep-learning setting. SAGE helps identify the species and data conditions for which a given approach is beneficial, thereby supporting the development of more transparent and ecologically credible SDMs. Data and code: https://earens.github.io/sage/
Ziang Fu, Ning Ning
Tue 29 Sep
cs.LGstat.ML
pdf html
- reading control, gains, good, state
- your work near “Timing as an Action: Learning When to Observe and Act” 0.71
- favorites near “Recent Developments in Machine Learning Methods for Stochastic Control and Games” 0.73
Abstract
Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly representations by directly straightening control geometry for sampling-efficient planning. CGS matches pairwise cosine similarities among actions to those among corresponding latent differences only using local transitions from pixel-action pairs. The loss can be applied across world-model architectures using end-to-end learned or pretrained representations. Under linear-dynamics, our theoretical analysis connects this objective to temporal straightening and more balanced terminal-cost curvature across the full planning horizon, yielding finite-budget guarantees for MPPI, local contraction results for CEM, and convergence bounds for gradient descent. Across four control environments and multiple planners, CGS improves planning with fewer sampled candidates and refinement steps, achieving success-rate gains up to 20 and 12.6 percentage points over LeWorldModel (LeWM) and its temporal-straightening variant (LeWM+TS), respectively, with sampling-based planners using 128 candidates per update. Probes, comparisons with DINO-WM architecture, and planner-side ablations clarify how latent motion organization, state dependence, and dynamical context shape planning behavior. Straightening control geometry thus makes good action sequences easier to find under limited planning budgets.
Shi-Shun Chen, Xiao-Yang Li
Mon 28 Sep
stat.AP
pdf html
- reading causal, processes, causal effects, wiener
- favorites near “Applied Koopmanism” 0.80
- your work near “Perturbation Methods for Incomplete Markets Economies” 0.68
Abstract
Accurate modeling of dependent degradation processes is essential for credible reliability assessment and remaining useful life (RUL) prediction in complex systems. Existing dependent degradation models usually describe dependence using correlation-based methods with symmetric characteristics. However, they ignore inherent causal directionality between degradation paths, which may lead to biased reliability evaluation and RUL predictions when physical causality exists. To address this issue, this paper proposes a causality-driven framework for modeling dependent degradation data. Firstly, univariate degradation models with multi-source uncertainties are established for each performance indicator based on the Wiener process. Then, the stable Peter-Clark algorithm is employed to uncover inherent causal relationships between degradation processes, and an uncertainty-aware neural network is employed to quantify causal effects considering uncertainties. Next, univariate degradation predictions and causal predictions are integrated within a Bayesian framework to construct a causally dependent degradation model, and the corresponding loss function is derived for model training. Finally, system reliability and RUL predictions are derived via Monte Carlo simulation. The proposed methodology is validated on the C-MAPSS dataset. Results show that considering inherent causal directionality between degradation processes helps eliminate physically unrealistic degradation behaviors, yielding more accurate degradation and RUL predictions than independent and correlation-based dependent degradation models.
Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang
Mon 28 Sep
cs.CRstat.AP
pdf html
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.81
- interest Language models as statistical objects 0.67
- your work near “Computational Methods for Economic Models with Function Valued States” 0.60
Abstract
Certified defenses that incorporate differential privacy have proven effective on Convolutional Neural Networks (CNNs), furnishing rigorous robustness guarantees against norm-bounded adversaries. However, the certified robustness behavior of Pixel Differential Privacy (PixelDP) remains largely unexplored with the self-attention architecture now dominating the deep-learning landscape. Given that the Transformer has a profound impact on our daily applications from the digital world to the physical world, it is crucial to study certified robustness through differential-privacy-style stability. To fill this research gap, we revisit this construction in Vision Transformers and identify a failure mode that is largely hidden in the convolutional setting. When noise is injected after the patch embedding, the Laplace mechanism with the inherited grouped $\ell_1$ sensitivity bound collapses to chance-level accuracy across noise scales, whereas the Gaussian mechanism remains trainable. This contrast isolates the source of failure: not the injected noise itself, but the geometry of the sensitivity constraint. We show that the attenuation induced by the inherited $\Delta_{1,1}$ projection increases with layer width and kernel size according to a random-matrix scale $C/(\sqrt{M}+\sqrt{N})$. Replacing the $\ell_1$-type constraint with a spectral-norm constraint eliminates the collapse across datasets and architectures, but creates a fundamental obstacle: the repaired models no longer satisfy the sensitivity condition required by the standard Laplace certificate. We resolve this mismatch by deriving a dimension-free $(\varepsilon,\ \delta)$-privacy guarantee for the Laplace mechanism under $\ell_2$ sensitivity through concentration of the privacy loss.
J. T. Korley, E. Adu Bonsu
Tue 29 Sep
stat.APstat.COstat.MEstat.OT
pdf html
- reading social, research, heterogeneity, national
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.69
- favorites near “Bayesian Semiparametric Model for Sequential Treatment Decisions with Informative Timing” 0.70
Abstract
Background. Magnetic resonance imaging (MRI) is central to Alzheimer's disease research and clinical evaluation, yet timing of first MRI after cohort entry varies substantially. In aging cohorts, many participants die before imaging, and treating death as noninformative censoring can distort inference on neuroimaging access and conceal disparities. Objective. To quantify individual-level determinants and center-level heterogeneity in timing of first neuroimaging while accounting for the competing risk of death. Methods. We analyzed National Alzheimer's Coordinating Center Uniform Data Set v3 linked to MRI records from 2015 through March 2025. The study included 20,867 participants aged 50-95 years without dementia at baseline and with known APOE {\epsilon}4 status. Time from baseline to first post-baseline MRI was the primary event, with death before MRI treated as a competing event. We fit Bayesian multilevel cause-specific accelerated failure time models with Alzheimer's Disease Center random intercepts, using Weibull and log-logistic models for MRI timing and death, respectively. Frequentist accelerated failure time and Cox models were used for sensitivity analyses. Results. APOE {\epsilon}4 was not credibly associated with timing of first MRI. Older age, lower educational attainment, and non-White race were associated with delayed MRI acquisition. Center-level variation in MRI timing was substantial. Conversely, APOE {\epsilon}4 homozygosity was associated with earlier death prior to MRI, alongside strong effects of age and cognitive impairment. Conclusions. Timing of first MRI in Alzheimer's research cohorts is driven primarily by demographic, social, and center-level factors rather than genetic risk. Accounting for competing mortality and site heterogeneity is necessary to avoid biased neuroimaging samples and support equitable Alzheimer's research.
Sikai Huang, Zhiwen Yang, Kai Yu, Jiayuan Chen, Stan Z. Li
Tue 29 Sep
cs.AIcs.LGq-bio.GNstat.AP
pdf html
- your work near “Valid Inference with Imperfect Synthetic Data” 0.72
- favorites near “Data Science at the Singularity” 0.79
- interest Language models as statistical objects 0.73
Abstract
Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.
Kangxiang Qin, Zeyu Li, Xinbing Kong, Wang Zhou
Mon 28 Sep
stat.ME
pdf html
- reading view, financial, recognition, slower
- interest Kernels, operators and functional data 0.72
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.72
Abstract
In this work, we model the observed multi-view tensors by decomposing the underlying signal in each view into two components: (i) the shared component that captures common dynamics across all views, and (ii) the private component that accounts for view-wise unique variations. To decouple the shared and private components, we introduce a novel Tucker personalized subspace principal component analysis (TPS-PCA) approach for tensors, which admits a one-step closed-form solution and serves as an ideal surrogate for our extended tensor-version personalized PCA (TP-PCA), adapted from the seminal work by \cite{shi2024personalized}. The theoretical analysis reveals that the proposed TPS-PCA estimators reach the minimax lower bound in terms of view-wise tensor decoupling, whereas the TP-PCA estimators only achieve a rate of average decoupling error across views, which is still slower than that of the TPS-PCA estimators. Extensive numerical experiments are conducted on synthetic and real datasets, demonstrating the wide applicability of the proposed method in fields including power management, financial analysis, and activity recognition.
Qizheng Xia, William Q. Meeker, Dennis Roach, Seth S. Kessler, Qing Li
Thu 24 Sep
stat.AP
pdf html
- reading structural, asset, technology, make
- favorites near “Applied Koopmanism” 0.78
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.67
Abstract
There is much interest in the potential to use structural health monitoring (SHM) technology to augment traditional nondestructive inspection (NDI) methods to improve safety, increase asset availability, and reduce maintenance and inspection costs. SHM has the potential to be used in many applications, including critical components in aircraft and pipelines. Probability of detection (POD) plays a critical role in aircraft structural integrity programs, leading to increased interest in developing methods to assess POD in SHM applications. In contrast to traditional NDI laboratory experiments involving specimens with cracks, SHM sensors are fixed, and SHM data are acquired over time as cracks grow or otherwise evolve. Thus, traditional statistical methods for assessing POD must be replaced or extended to properly handle repeated-measures data. The purpose of this paper is to review the basic statistical concepts of POD and show how these concepts can be extended or adapted for SHM-POD applications. The paper presents statistical methods for modifying and extending existing POD methods, including a simple size-of-damage-at-detection (SoDaD) method and a random-parameter (RP) method for repeated-measures data. The methods are compared using three case studies involving Piezoelectric Transducer (PZT), Carbon Nanotube (CNT), and Comparative Vacuum Monitoring (CVM) sensor systems. Results show that the SoDaD method provides a simple approach for POD estimation with limited data, while the RP method offers enhanced modeling fidelity by utilizing repeated measurements. These methods are applicable when a scalar damage index or similar response is used to make a detection decision.
Noriaki Hashimoto, Shuichi Nishino, Teruyuki Katsuoka, Tomohiro Shiraishi, Daiki Miwa, Hiroyuki Hanada, et al.
Tue 29 Sep
cs.CVcs.LGstat.ML
pdf html
- favorites near “Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss” 0.78
- interest Kernels, operators and functional data 0.70
- reading procedures, inference, higher, control
Abstract
Multiple instance learning (MIL) is widely used in computational pathology because it enables weakly supervised analysis of whole-slide images (WSIs) without requiring patch-level annotations. In attention-based MIL, instances with high attention scores are often interpreted as diagnostically important regions and used as visual explanations. However, attention scores alone cannot determine whether selected high-attention instances are significantly different from normal instances, limiting the reliability of attention-based explanations. In this paper, we formulate the evaluation of high-attention instances as a statistical hypothesis testing problem. Specifically, we assess whether a selected high-attention instance significantly deviates from a representative normal reference instance selected based on feature similarity. A major challenge is that both the target instance and the reference instance are selected through data-dependent procedures, rendering standard hypothesis testing invalid. To address this issue, we introduce a selective inference (SI) framework that explicitly accounts for the selection events induced by attention-based instance selection and adaptive reference selection, thereby enabling the computation of valid selective $p$-values conditional on these events. Experiments demonstrate Type-I error control on synthetic and MNIST-based data and practical applicability to pathological WSIs, with higher statistical power than the conventional over-conditioning approach.
Enock Adu Bonsu
Wed 30 Sep
stat.ME
pdf html
- your work near “Local Causal Discovery for Estimating Causal Effects” 0.76
- reading did, procedures, observable, discrimination
- interest Empirical Bayes, shrinkage and decision theory 0.76
- looks like clinical and biomedical studies
Abstract
AI models trained on the same data can disagree about patient risk, with disagreement potentially concentrated in clinically important subgroups. We propose V(S), a statistically grounded vulnerability index combining an observable lower-bound witness of model disagreement with clinical severity, and develop inference and multiplicity-adjustment procedures for auditing prespecified subgroups. We applied the framework to two large critical-care cohorts, MIMIC-IV (n = 65,078) for model development and eICU-CRD (n = 188,230 admissions, 208 hospitals) for external validation, comparing a random forest with logistic regression across 158 prespecified subgroups. The two primary models did not both satisfy the prespecified epsilon = 0.02 Rashomon-set tolerance: the logistic-regression AUC was 0.0488 below the best candidate-model AUC. The RF-LR discrimination gap is thus interpreted as disagreement between two specific models, not as a guaranteed lower bound on the full Rashomon set. Nine subgroups had discrimination gaps distinguishable from a prespecified clinical floor. The age >=80 and cardiac subgroup had the largest point estimate of V(S) (0.307), but was underpowered and did not meet the full high-priority decision rule. The univariate cardiac subgroup (V(S) = 0.193, 95% CI [0.163, 0.223]) was the only statistically distinguishable subgroup with adequate power. Post hoc analyses identified lactate as important for both models but did not establish a causal explanation for the disagreement. Four simulation studies quantified operating characteristics of the proposed procedures, including inflated small-sample detection rates and imperfect Wald-interval coverage. The framework offers a reproducible approach for ranking subgroup vulnerability to model disagreement while separating exploratory signals from adequately supported findings.
Cao Le Cong Thanh, Dang Quang Vinh, Vo Nguyen Le Duy
Wed 30 Sep
stat.MLcs.LG
pdf html
- interest Kernels, operators and functional data 0.73
- favorites near “Why Does Deep Learning Work? - A Perspective From Group Theory” 0.74
- your work near “Online Data Collection for Efficient Semiparametric Inference” 0.69
Abstract
Deep Support Vector Data Description (Deep SVDD) has become a prominent framework for unsupervised anomaly detection by learning latent representations that compactly characterize normal data around a center. Despite its empirical success, anomaly decisions produced by Deep SVDD are typically made solely based on anomaly scores without rigorous statistical guarantees, thereby limiting their reliability in safety-critical and high-stakes applications where false positives must be strictly controlled. In this paper, we propose PADI (Post-Anomaly Detection Inference), a novel framework that equips a trained and frozen Deep SVDD detector with statistically valid inference by leveraging the Selective Inference framework. Specifically, PADI performs inference conditional on the event that a test instance is identified as anomalous by Deep SVDD, thereby enabling rigorous statistical assessment of anomaly decisions. Based on this formulation, we derive valid selective p-values that quantify the statistical significance of the detected anomaly. Using these p-values, we theoretically establish control of the false positive rate (FPR) at a user-specified significance level $\alpha$ (e.g., $\alpha=0.05$). Furthermore, we extend the proposed framework to Deep Semi-Supervised Anomaly Detection (Deep SAD), providing a principled approach for statistically reliable inference in semi-supervised anomaly detection settings. Extensive experiments on both synthetic and real-world benchmark datasets robustly support the theoretical findings. The results demonstrate that PADI consistently achieves proper FPR control while attaining superior true positive rates compared with existing approaches.
Haixiang Sun, Andrew L. Liu
Mon 28 Sep
stat.MLcs.LGmath.OC
pdf html
- interest Kernels, operators and functional data 0.73
- your work near “Timing as an Action: Learning When to Observe and Act” 0.68
- favorites near “Monte Carlo Geometry Processing” 0.74
Abstract
Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases that finite datasets fail to capture, making simple resampling or perturbation insufficient for stress scenario generation. Existing outlier synthesis methods typically rely on sparse neighborhoods, low support latent regions, or classifier boundary crossings, which can be heuristic, unstable, and tied to specific modalities or architectures. We therefore propose Sinkhorn Boundary Outlier Generation (SBOG), a structured framework for latent-space outlier generation that couples Sinkhorn optimal transport geometry with distributionally robust boundary modeling. The resulting Sinkhorn-induced support cost guides the sampler toward weakly supported boundary regions, while semantic constraints prevent uncontrolled drift from the intended context, yielding controlled deviations from the in-distribution reference measure rather than arbitrary sparse-region samples. Experiments on time series anomaly generation and image outlier synthesis show that our framework produces informative, semantically controlled outliers and improves downstream robustness evaluation across modalities, providing a foundation for stress scenario generation beyond empirical support.
Jaume Aguiló Piña, Laia Subirats, Aina Frau-Pascual, Alba Gorriz, Rudys Magrans Nicieza, Jordi Morillas Perez
Tue 29 Sep
stat.AP
pdf html
- reading shock, units, early, causes
- interest Language models as statistical objects 0.68
- your work near “Computational Methods for Economic Models with Function Valued States” 0.61
- looks like domain applications of machine learning
Abstract
Circulatory shock is one of the leading causes of mortality in intensive care units (ICUs), and its early detection is critical to enable timely treatment and improve clinical outcomes. This study aimed to develop and evaluate a two-stage cascade machine learning framework for the early detection and etiological classification of circulatory shock in critically ill patients. Using data from the Medical Information Mart for Intensive Care (MIMIC)-IV database, four patient groups were defined: septic shock, cardiogenic shock, hypovolemic shock, and a non?shock control group, comprising a total of 32,907 patients. Vital signs and laboratory data were collected during the first six hours after ICU admission. After data cleaning and missing-value imputation, the mean value of each variable was used for model development. Several machine learning algorithms were compared, including logistic regression, Random Forest, XGBoost, and multilayer perceptron (MLP) networks. Random Forest and XGBoost achieved the highest overall performance, with an AUROC of approximately 0.82-0.83 for shock detection, and a macro-averaged sensitivity of approximately 0.61 and precision of approximately 0.58 across all four classes. Classification performance was highest for the non?shock group, followed by septic and cardiogenic shock, while hypovolemic shock showed the lowest performance. These results indicate that machine learning models can identify early signs of hemodynamic deterioration associated with circulatory shock and may support clinical decision-making in the ICU. However, further improvements are needed for the classification of specific shock subtypes, particularly hypovolemic and cardiogenic shock, as well as for real-time clinical implementation.