K-means cluster analysis, Advanced Statistics

Assignment Help:

K-means cluster analysis is the method of cluster analysis in which from an initial partition of observations into K clusters, each observation in turn is analysed and reassigned, if suitable, to a different cluster in an attempt to optimize some predefined numerical criterion that measures in some sense the 'quality' of cluster solution. Several such clustering criteria have been suggested, but the most usually used arise from considering the features of the within groups, between groups and whole matrices of sums of squares and the cross products (W, B, T) which can be described for every partition of the observations into the particular number of groups. The two most ordinary of the clustering criteria developing from these matrices are given as follows

minimization of trace W

minimization of determinant W

The first of these has tendency to produce the 'spherical' clusters, the second to produce clusters that all have same shape, though this will not necessarily be spherical in shape. 

 


Related Discussions:- K-means cluster analysis

Higher criticism, Higher criticism is a multiple-comparison test concept a...

Higher criticism is a multiple-comparison test concept arising from the situation where there are number of independent tests of significance and interest lies in the rejecting jo

Procrustes analysis, Procrustes analysis is a technique of comparing the a...

Procrustes analysis is a technique of comparing the alternative geometrical representations of a group of multivariate data or of the proximity matrix, for instance, two competing

Explain influence statistics, Influence statistics: The range of statistic...

Influence statistics: The range of statistics designed to assess the effect or the in?uence of an observation in determining results of the regression analysis. The general approa

Current status data, The Current status data arise in the survival analysis...

The Current status data arise in the survival analysis if the observations are limited to the indicators of whether or not the event of interest has happened at the time the sample

Non central distributions, Non central distributions is the series of prob...

Non central distributions is the series of probability distributions each of which is the adaptation of one of the standard sampling distributions like the chi-squared distributio

Hirap, #q A paper mill products two grade of paper viz., X & Y. Because of ...

#q A paper mill products two grade of paper viz., X & Y. Because of raw material restriction, it cannot produce more than 400 tons of grade X paper & 300 tons of grade Y paper in a

Pattern recognition, Pattern recognition is a term for a technology that r...

Pattern recognition is a term for a technology that recognizes and analyses patterns automatically by machine and which has been used successfully in many areas of application inc

Helmert contrast, Helmert contrast is the contrast often used in analysis ...

Helmert contrast is the contrast often used in analysis of the variance, in which each level of a factor is tested against average of the remaining levels. So, for instance, if th

Expectaton, sales per day for a product are as follows: x= 10, 11, 12, 13 (...

sales per day for a product are as follows: x= 10, 11, 12, 13 (p)= 0.2, 0.4, 0.3, 0.1 obtain mean and variance of daily sale. if the profit is described by the following equation p

Factorial designs, Designs which permits two or more questions to be addres...

Designs which permits two or more questions to be addressed in the investigation. The easiest factorial design is one in which each of the two treatments or interventions are p

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd