|
In statistical theory, a U-statistic is a class of statistics that is especially important in estimation theory; the letter "U" stands for unbiased. In elementary statistics, U-statistics arise naturally in producing minimum-variance unbiased estimators. The theory of U-statistics allows a minimum-variance unbiased estimator to be derived from each unbiased estimator of an ''estimable parameter'' (alternatively, ''statistical functional'') for large classes of probability distributions.〔Cox & Hinkley (1974),p. 200, p. 258〕〔Hoeffding (1948), between Eq's(4.3),(4.4)〕 An estimable parameter is a measurable function of the population's cumulative probability distribution: For example, for every probability distribution, the population median is an estimable parameter. The theory of U-statistics applies to general classes of probability distributions. Many statistics originally derived for particular parametric families have been recognized as U-statistics for general distributions. In non-parametric statistics, the theory of U-statistics is used to establish for statistical procedures (such as estimators and tests) and estimators relating to the asymptotic normality and to the variance (in finite samples) of such quantities.〔Sen (1992)〕 The theory has been used to study more general statistics as well as stochastic processes, such as random graphs.〔Page 508 in 〕〔Pages 381–382 in 〕〔Page xii in 〕 Suppose that a problem involves independent and identically-distributed random variables and that estimation of a certain parameter is required. Suppose that a simple unbiased estimate can be constructed based on only a few observations: this defines the basic estimator based on a given number of observations. For example, a single observation is itself an unbiased estimate of the mean and a pair of observations can be used to derive an unbiased estimate of the variance. The U-statistic based on this estimator is defined as the average (across all combinatorial selections of the given size from the full set of observations) of the basic estimator applied to the sub-samples. Sen (1992) provides a review of the paper by Wassily Hoeffding (1948), which introduced U-statistics and set out the theory relating to them, and in doing so Sen outlines the importance U-statistics have in statistical theory. Sen says〔Sen (1992) p. 307〕 "The impact of Hoeffding (1948) is overwhelming at the present time and is very likely to continue in the years to come". Note that the theory of U-statistics is not limited to〔Sen (1992), p306〕 the case of independent and identically-distributed random variables or to scalar random-variables.〔Borovskikh's last chapter discusses U-statistics for exchangeable random elements taking values in a vector space (separable Banach space).〕 ==Definition== The term U-statistic, due to Hoeffding (1948), is defined as follows. Let be a real-valued or complex-valued function of variables. For each the associated U-statistic is equal to the average over ordered samples of size of the sample values . In other words, , the average being taken over distinct ordered samples of size taken from . Each U-statistic is necessarily a symmetric function. U-statistics are very natural in statistical work, particularly in Hoeffding's context of independent and identically-distributed random variables, or more generally for exchangeable sequences, such as in simple random sampling from a finite population, where the defining property is termed `inheritance on the average'. Fisher's ''k''-statistics and Tukey's polykays are examples of homogeneous polynomial U-statistics (Fisher, 1929; Tukey, 1950). For a simple random sample ''φ'' of size ''n'' taken from a population of size ''N'', the U-statistic has the property that the average over sample values ''ƒ''''n''(''xφ'') is exactly equal to the population value ''ƒ''''N''(''x''). 抄文引用元・出典: フリー百科事典『 ウィキペディア(Wikipedia)』 ■ウィキペディアで「U-statistic」の詳細全文を読む スポンサード リンク
|