K-means cluster analysis, Advanced Statistics

Assignment Help:

K-means cluster analysis is the method of cluster analysis in which from an initial partition of observations into K clusters, each observation in turn is analysed and reassigned, if suitable, to a different cluster in an attempt to optimize some predefined numerical criterion that measures in some sense the 'quality' of cluster solution. Several such clustering criteria have been suggested, but the most usually used arise from considering the features of the within groups, between groups and whole matrices of sums of squares and the cross products (W, B, T) which can be described for every partition of the observations into the particular number of groups. The two most ordinary of the clustering criteria developing from these matrices are given as follows

minimization of trace W

minimization of determinant W

The first of these has tendency to produce the 'spherical' clusters, the second to produce clusters that all have same shape, though this will not necessarily be spherical in shape. 

 


Related Discussions:- K-means cluster analysis

Finite population correction, This term sometimes used to describe the extr...

This term sometimes used to describe the extra factor in variance of the sample mean when n sample values are drawn without the replacement from the finite population of size N. Th

Generalized additive model, The linear component ηi, de?ned just in the tra...

The linear component ηi, de?ned just in the traditional way: η i = x' 1 A monotone differentiable link function g that describes how E(Yi) = µi is related to the linear compon

Link functions, Link functions: The link function relates the linear p...

Link functions: The link function relates the linear predictor ηi to the expected value of the data. In classical linear models the mean and the linear predictor are identical

Weathervane plot, Weathervane plot is the graphical display of the multiva...

Weathervane plot is the graphical display of the multivariate data based on bubble plot. The latter is enhanced by the addiction of the lines whose lengths and directions code the

Randomization tests, Randomization tests are the procedures for determinin...

Randomization tests are the procedures for determining the statistical significance directly from the data with- out recourse to some particular sampling distribution. For instanc

Genomics, Genomics  is the study of the structure, function and the evoluti...

Genomics  is the study of the structure, function and the evolution of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sequences which comprise the genome of living organisms

Likelihood, Likelihood is the probability of a set of observations provide...

Likelihood is the probability of a set of observations provided the value of some parameter or the set of parameters. For instance, the likelihood of the random sample of n observ

Math, A standard IQ test has a mean of 98 and a standard deviation of 16. W...

A standard IQ test has a mean of 98 and a standard deviation of 16. We want to be 99% certain that we are within 8 IQ points of the true mean. Determine the sample size

Disease mapping, The method of displaying the geographical variability of t...

The method of displaying the geographical variability of the disease on maps using different colors, shading, etc. The logic is not new, but the arrival of computers and computer g

Descriptive , Assume that a population is normally distributed with a mean ...

Assume that a population is normally distributed with a mean of 100 and a standard deviation of 15. Would it be unusual for the mean of a sample of 20 to be 115 or more?

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd