Transformation of data, Applied Statistics

Assignment Help:

PCA is a linear transformation that transforms the data to a new coordinate system such that the greatest variance by any projection of the data comes to lie on the first coordinate (called the first principal component), the second greatest variance on the second coordinate, and so on. The PCA can be used for dimensionality reduction in a dataset while retaining those characteristics of the dataset that contribute most to its variance, by keeping lower-order principal components and ignoring higher-order ones. Such low-order components often contain the "most important" aspects of the data. But this is not necessarily the case, depending on the application. Let p and tn denote respectively the original and reduced number of variables. The original variables are denoted X. In the simplest case our measure of accuracy of reconstruction is the sum ofp squared multiple correlations between X-variables and the predictions of X made froin the factors. In the more general case we can weight each squared multiple correlation by the variance of the corresponding X-variable.

Since we can set those variances ourselves by multiplying scores on each variable,by any constant we choose, this amounts to the ability to assign any weights we choose to the different variables.


Related Discussions:- Transformation of data

Frequency distribution, mark number of student 0-10 4 10-20 8 ...

mark number of student 0-10 4 10-20 8 20-30 11 30-40 15 40-50 12 50-60 6 calculate frequency distribution

Calculate the one year interest rate, A.The coupon rate of Erie-Chicago Rai...

A.The coupon rate of Erie-Chicago Rail is 7%. The interest rate of Florida municipal bond with equal risk is 6%.  At what tax rate the two bonds are as good as each other B.Supp

Prediction, Differentiate between prediction, projection and forecasting.

Differentiate between prediction, projection and forecasting.

Statistical procedures - estimation of a mean, Old Faithful Geyser in Yello...

Old Faithful Geyser in Yellowstone National Park derives its names and fame from the regularity (and beauty) of its eruptions. Rangers usually post the predicted times of eruptions

Canonical correlation analysis, Canonical correlation analysis (CC) allows ...

Canonical correlation analysis (CC) allows the investigation of the relationship between two ,sets of variables. For example, a sociologist may want to investigate the Relationship

Residual, regression line drawn as Y=C+1075x, when x was 2, and y was 239, ...

regression line drawn as Y=C+1075x, when x was 2, and y was 239, given that y intercept was 11. calculate the residual

Schedule, Schedule Schedule is also used for the collection of primary ...

Schedule Schedule is also used for the collection of primary data. A schedule is a list of question. it is a device of obtaining answer to the questions in a form which is fill

Vital statistics, How vital statistics are affects on our human life

How vital statistics are affects on our human life

..National Account- Descriptive Statistics, A country''s national accounts ...

A country''s national accounts are assumed to look as follows: GDP 1180 VAT and taxes 140 Commodity subsidies 60 Raw material and consumables 530 1. Calculate GVA 2. Calculate t

Eliminate all of the insignificant variables, The file Midterm Data.xls ha...

The file Midterm Data.xls has a tab labeled "Many vs. S&P" which presents historical price data for several assets, a volatility condition (VIDX = 1 if the NYSE volatility is grea

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd