Outliers - reasons for screening data, Advanced Statistics

Assignment Help:

Outliers - Reasons for Screening Data

Outliers are due to data entry errors, subject is not a member of the population that the sample is trying to represent, or the subject is really different. Statistical tests are quite sensitive to outliers so this problem should be addressed.

Univariate outliers are easy to detect (z-scores, box plots, histograms, etc.) standard scores larger than +/-3 are outliers (consider 4 is n>100 or 2.5 if n<10)

Multivariate outliers are difficult to detect. Mahalanobis distance is one powerful technique to use in this case (discussed later). This is evaluated as a chi-square statistic with degrees of freedom equal to number of variables in the analysis. A chi-sqaure statistic value that is significant beyond p<0.001 level determines outliers.

In most cases, it is ok to drop the value from the sample. One can also take steps to reduce the relative influence of outliers if the researcher decides to include the values in the analysis.


Related Discussions:- Outliers - reasons for screening data

Variance inflation factor, VIF is the abbreviation of variance inflation fa...

VIF is the abbreviation of variance inflation factor which is a measure of the amount of multicollinearity that exists in a set of multiple regression variables. *The VIF value

Infant mortality rate, Infant mortality rate is the ratio of the number of...

Infant mortality rate is the ratio of the number of deaths during the calendar year among the infants under one year of age to the total number of live births during that particul

Current status data, The Current status data arise in the survival analysis...

The Current status data arise in the survival analysis if the observations are limited to the indicators of whether or not the event of interest has happened at the time the sample

Linear regression assignment help, Using World Bank (2004) World Developmen...

Using World Bank (2004) World Development Indicators; Washington: International Bank for Reconstruction & Development/ The World Bank, located in the reference section of the Learn

Mean-range plot, Mean-range plot   is the graphical tool or device usefu...

Mean-range plot   is the graphical tool or device useful in selecting a transformation in the time series analysis. The range is plotted against the mean for each of the seasona

Explain post stratification adjustment, Post stratification adjustmen t: On...

Post stratification adjustmen t: One of the most often used population weighting adjustments used in the complex surveys, in which weights for the elements in a class are multiplie

Game theory, This is the branch of mathematics which deals with the theory ...

This is the branch of mathematics which deals with the theory of contests between two or more players under the specified sets of rules. The subject supposes a statistical aspect w

F-test, A test for equality of the variances of the two populations having ...

A test for equality of the variances of the two populations having normal distributions, based on the ratio of the variances of the sample of observations taken from each. Most fre

Explanatory variables, The variables appearing on the right-hand side of eq...

The variables appearing on the right-hand side of equations defining, for instance, multiple regressions or the logistic regression, and which seek to predict or 'explain' response

Matching, Matching is the method of making a study group and a comparison ...

Matching is the method of making a study group and a comparison group comparable with respect to the extraneous factors. Generally used in the retrospective studies when selecting

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd