K-means cluster analysis, Advanced Statistics

Assignment Help:

K-means cluster analysis is the method of cluster analysis in which from an initial partition of observations into K clusters, each observation in turn is analysed and reassigned, if suitable, to a different cluster in an attempt to optimize some predefined numerical criterion that measures in some sense the 'quality' of cluster solution. Several such clustering criteria have been suggested, but the most usually used arise from considering the features of the within groups, between groups and whole matrices of sums of squares and the cross products (W, B, T) which can be described for every partition of the observations into the particular number of groups. The two most ordinary of the clustering criteria developing from these matrices are given as follows

minimization of trace W

minimization of determinant W

The first of these has tendency to produce the 'spherical' clusters, the second to produce clusters that all have same shape, though this will not necessarily be spherical in shape. 

 


Related Discussions:- K-means cluster analysis

Decision Models., An oil company thinks that there is a 60% chance that the...

An oil company thinks that there is a 60% chance that there is oil in the land they own. Before drilling they run a soil test. When there is oil in the ground, the soil test comes

Distance sampling, The technique of sampling used in the ecology for determ...

The technique of sampling used in the ecology for determining how much plants or animals are in a given fixed region. A set of randomly placed lines or points is recognized and the

Regression discontinuity design, Regression discontinuity design is the qu...

Regression discontinuity design is the quasi-experimental design in which participants in, for instance, an intervention study, are assigned to the treatment and control groups on

Determine the maximum amount of the commodity, A manufacturing company has ...

A manufacturing company has two factories F 1 and F 2 producing a certain commodity that is required at three retail outlets M 1 , M 2 and M 3 . Once produced, the commodity is

Incubation period, Incubation period is the time elapsing amongs the receip...

Incubation period is the time elapsing amongs the receipt of infection and the appearance of the symptoms. The length of the incubation time period depends on the disease, ranging

Hosmer-lemeshow test, Hosmer-Lemeshow test is a goodness-of-fit test taken...

Hosmer-Lemeshow test is a goodness-of-fit test taken in use in logistic regression, particularly when there are regular covariates. Units are spitted into deciles based on predict

Morbidity, Morbidity is the term used in the epidemiological studies to de...

Morbidity is the term used in the epidemiological studies to describe sickness in the human populations. The WHO Expert Committee on the Health Statistics noted in its sixth repor

Randomized encouragement trial, Randomized encouragement trial   is the cl...

Randomized encouragement trial   is the clinical trials in which the participants are encouraged to change their behaviour in a particular manner (or not, if they are allocated to

Balanced incomplete block design, Balanced incomplete block design : A desi...

Balanced incomplete block design : A design in which all the treatments are not used in all blocks. Such designs have the below stated properties: * each block comprises the

Explain johnson-neyman technique, Johnson-Neyman technique:  The technique ...

Johnson-Neyman technique:  The technique which can be used in the situations where analysis of the covariance is not valid because of the heterogeneity of slopes. With this method

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd