What is the effect of lowering the number

Assignment Help Other Subject
Reference no: EM132079366

PART 1: CLASSIFICATION

This part of the assignment is concerned with the file:

/KDrive/SEH/SCSIT/Students/Courses/COSC2111/DataMining/ data/other/bank-balanced1.csv.

There is a description of the data in the file bank-names.txt in the same directory. [bank-balanced1.csv is a subset of bank-full.csv]. The main goal is to achieve the highest classification accuracy with the lowest amount of overfitting.

1. Run the following classifiers, with the default parameters, on this data: ZeroR, OneR, J48, IBK and construct a table of the training and cross-validation errors. You can get the training error by selecting "Use training set" as the test option. What do you conclude from these results?

Run No

Classifier

Parameters

Parameters

Training

Error

Cross-valid

Error

Over-

Fitting

1

.

ZeroR

.

None

.

30.0%

.

30.0%

.

None

2. Using the J48 classifier, can you find a combination of the C and M parameter values that minimizes the amount of overfitting? Include the results of your best five runs, including the parameter values, in your table of results.

3. Reset J48 parameters to their default values. What is the effect of lowering the number of examples in the training set? Include your runs in your table of re- sults.

4. Using the IBk classifier, can you find the value of k that minimizes the amount of overfitting? Include your runs in your table of results.

5. Try a number of other classifiers. Aside from ZeroR, which classifiers are best and worst in terms of predictive accuracy? Include 5 runs in your table of results.

6. Compare the accuracy of ZeroR, OneR and J48. What do you conclude?

7. What golden nuggets did you find, if any?

8. [OPTIONAL] Use an attribute selection algorithm to get a reduced attribute set. How does the accuracy on the reduced set compare with the accuracy on the full set?

Report Length: Up to two pages.

PART 2: NUMERIC PREDICTION

Numeric Prediction of the balance attribute in the bank data of part 1. The main goal is to achieve the lowest mean absolute error with the lowest amount of overfitting.

1. Run the following classifers, with default parameters, on this data: ZeroR, MP5, IBk and construct a table of the training and cross-validation errors. You may want to turn on "Output Predictions" to get a better sense of the magnitude of the error on each example. What do you conclude from these results?

2. Explore different parameter settings for M5P and IBk. Which values give the best performance in terms of predictive accuracy and overfitting. Include the results of the best five runs in your table of results.

3. Investigate three other classifiers for numeric prediction and their associated pa- rameters. Include your best five runs in your table of results. Which classifier gives the best performance in terms of predictive accuracy and overfitting?

4. What golden nuggets did you find, if any?

Report Length Up to one page.

PART 3: CLUSTERING

Clustering of the bank data of part 1. For this part use only the attributes age, marital, education, and balance.

The aim is determine the number of clusters in the data and assess whether any of the clusters are meaningful.

1. Run the Kmeans clustering algorithm on this data for the following values of K: 1,2,3,4,5,10,20. Analyse the resulting clusters. What do you conclude?

2. Choose a value of K and run the algorithm with different seeds. What is the effect of changing the seed?

3. Run the EM algorithm on this data with the default parameters and describe the output.

4. The EM algorithm can be quite sensitive to whether the data is normalized or not. Use the weka normalize filter (Preprocess --> Filter --> unsupervised --> normalize) to normalize the numeric attributes. What difference does this make to the clus- tering runs?

5. The algorithm can be quite sensitive to the values of minLogLikelihoodImprove- mentCV minStdDev and minLogLikelihoodImprovementIterating, Explore the effect of changing these values. What do you conclude?

6. How many clusters do you think are in the data? Give an English language description of one of them.

7. Compare the use of Kmeans and EM for these clustering tasks. Which do you think is best? Why?

8. What golden nuggets did you find, if any?
Report Length Up to one page.

PART 4: ASSOCIATION FINDING

Association finding in the files supermarket1.arff and supermarket2.arff in the folder
/KDrive/SEH/SCSIT/Students/Courses/COSC2111/DataMining/data/arff.

The main aim is to determine whether there are any significant associations in the data.

These files contain the same details of shopping transactions represented in two different ways. You can use a text viewer to look at the files.

1. What is the difference in representations?

2. Load the file supermarket1.arff into weka and run the Apriori algorithm on this data. You might need to restrict the number of attributes and/or the number of examples. What significant associations can you find?

3. Explore different possibilities of the metric type and associated parameters. What do you find?

4. Load the file supermarket22.arff into weka and run the Apriori algorithm on this data. What do you find?

5. Explore different possibilities of the metric type and associated parameters. What do you find?

6. Try the other associators. What are the differences to Apriori?

7. What golden nuggets did you find, if any?

8. [OPTIONAL] Can you find any meaningful associations in the bank data?

Report Length Up to one page.

Verified Expert

The solution file resolved the laboratory practical problems on data mining using data obtained from University K drive. The software used for data mining problem is Weka-3-9-2 obtained from RMIT University. The results were saved in the same access.

Reference no: EM132079366

Questions Cloud

Describe the type of diversification : Describe the type(s) of diversification that Lincoln Electric has pursued. Be specific with examples.
Reports or presentations to other managers and clients : Can anyone clarify some ways that businesses can conserve resources, especially when having to provide reports or presentations to other managers and clients.
Launching a new product line : When it comes to starting a business or launching a new product line, most people do not apply much (if any) scrutiny to their business licenses
Why you have chosen this topic for your literature review : Write one to two paragraphs (a) summarizing the problem area (be specific in defining the problem), (b) describing what you already know about the topic.
What is the effect of lowering the number : COSC2110/COSC2111 Data Mining - RMIT University - What is the effect of lowering the number of examples in the training set? Include your runs in your table
What is this question intended to address : Begin the review by defining the objective of the paper. Introduce the reader to your focal question. What is this question intended to address?
Do companies want us to believe that they are invested : Do companies want us to believe that they are invested in CSR? How can we tell?
Social responsibility and commitment to stakeholders : In studying corporate social responsibility and commitment to stakeholders, what obligation does a company have to it's stakeholders?
Identify problem areas and how they impact the organization : Identify 2-3 problem areas and how they impact the organization. Consider 2-4 solutions through course readings, research, and your experience.

Reviews

inf2079366

10/31/2018 3:41:52 AM

In this assignment, you are asked to apply a number of algorithms to a number of data sets and write a report on your findings. You will be assessed on methodology, analysis of results and conclusions. for the following given assignment, I have received a proper solution ExpertsMind are a reliable source for students who have limited time to do an assignment on their own. It is a wonderful organisation that interested in students success. I really thankful to them.

len2079366

8/7/2018 10:28:51 PM

This assignment counts for 23% of the total marks in this course. Due Date 9:00am Monday 27 Submit through Canvas You can work on this assignment individually or in a group of 2. In this assignment you are asked to apply a number of algorithms to a number of data sets and write a report on your findings. You will be assessed on methodology, analysis of results and conclusions.

Write a Review

Other Subject Questions & Answers

  Cross-cultural opportunities and conflicts in canada

Short Paper on Cross-cultural Opportunities and Conflicts in Canada.

  Sociology theory questions

Sociology are very fundamental in nature. Role strain and role constraint speak about the duties and responsibilities of the roles of people in society or in a group. A short theory about Darwin and Moths is also answered.

  A book review on unfaithful angels

This review will help the reader understand the social work profession through different concepts giving the glimpse of why the social work profession might have drifted away from its original purpose of serving the poor.

  Disorder paper: schizophrenia

Schizophrenia does not really have just one single cause. It is a possibility that this disorder could be inherited but not all doctors are sure.

  Individual assignment: two models handout and rubric

Individual Assignment : Two Models Handout and Rubric,    This paper will allow you to understand and evaluate two vastly different organizational models and to effectively communicate their differences.

  Developing strategic intent for toyota

The following report includes the description about the organization, its strategies, industry analysis in which it operates and its position in the industry.

  Gasoline powered passenger vehicles

In this study, we examine how gasoline price volatility and income of the consumers impacts consumer's demand for gasoline.

  An aspect of poverty in canada

Economics thesis undergrad 4th year paper to write. it should be about 22 pages in length, literature review, economic analysis and then data or cost benefit analysis.

  Ngn customer satisfaction qos indicator for 3g services

The paper aims to highlight the global trends in countries and regions where 3G has already been introduced and propose an implementation plan to the telecom operators of developing countries.

  Prepare a power point presentation

Prepare the power point presentation for the case: Santa Fe Independent School District

  Information literacy is important in this environment

Information literacy is critically important in this contemporary environment

  Associative property of multiplication

Write a definition for associative property of multiplication.

Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd