What is the overall error for the validation set

Assignment Help Computer Engineering
Reference no: EM131925054

Problem

Automobile Accidents. The file Accidents.csv contains information on 42,183 actual automobile accidents in 2001 in the United States that involved one of three levels of injury: NO INJURY, INJURY, or FATALITY. For each accident, additional information is recorded, such as day of week, weather conditions, and road type. A firm might be interested in developing a system for quickly classifying the severity of an accident based on initial reports and associated data in the system (some of which rely on GPS-assisted reporting).

Our goal here is to predict whether an accident just reported will involve an injury (MAX_SEV_IR = 1 or 2) or will not (MAX_SEV_IR = 0). For this purpose, create a dummy variable called INJURY that takes the value "yes" if MAX_SEV_IR = 1 or 2, and otherwise "no."

a. Using the information in this dataset, if an accident has just been reported and no further information is available, what should the prediction be? (INJURY = Yes or No?) Why?

b. Select the first 12 records in the dataset and look only at the response (INJURY) and the two predictors WEATHER_R and TRAF_CON_R.

i. Create a pivot table that examines INJURY as a function of the two predictors for these 12 records. Use all three variables in the pivot table as rows/columns.

ii. Compute the exact Bayes conditional probabilities of an injury (INJURY = Yes) given the six possible combinations of the predictors.

iii. Classify the 12 accidents using these probabilities and a cutoff of 0.5.

iv. Compute manually the naive Bayes conditional probability of an injury given WEATHER_R = 1 and TRAF_CON_R = 1.

v. Run a naive Bayes classifier on the 12 records and two predictors using R. Check the model output to obtain probabilities and classifications for all 12 records. Compare this to the exact Bayes classification. Are the resulting classifications equivalent? Is the ranking (= ordering) of observations equivalent?

c. Let us now return to the entire dataset. Partition the data into training (60%) and validation (40%).

i. Assuming that no information or initial reports about the accident itself are available at the time of prediction (only location characteristics, weather conditions, etc.), which predictors can we include in the analysis? (Use the _ sheet.)

ii. Run a naive Bayes classifier on the complete training set with the relevant predictors (and INJURY as the response). Note that all predictors are categorical. Show the confusion matrix.

iii. What is the overall error for the validation set?

iv. What is the percent improvement relative to the naive rule (using the validation set)?

v. Examine the conditional probabilities output. Why do we get a probability of zero for P(INJURY = No | SPD_LIM = 5)?

Reference no: EM131925054

Questions Cloud

What is the difference between pure and mixed ip models : Provide your own examples of five applications of IP. What is the difference between pure and mixed IP models? Which do you think is most common, and why?
Calculate the variance of portfolio returns : Calculate the variance of portfolio returns, assuming the correlation between the returns is 1.
Write new server and client programs and probe and sampler : Write new server and client programs, Probe and Sampler, that use our Server and Client classes, respectively.
Calculate the new beps if it were possible to raise : The text shows how to compute the break-even point in units; sometimes it's also useful to know the break-even point in dollars (sales).
What is the overall error for the validation set : What is the overall error for the validation set? What is the percent improvement relative to the naive rule (using the validation set)?
Find the weighted average cost of this capital : Company A agrees to lend $300,000 and they require 5% interest, Company B will lend $200,000 at 6% interest, and Company C will loan the balance.
Why do you expect to see seasonality in sales of shampoo : Why do you expect to see seasonality in sales of shampoo? Why? If the goal is forecasting sales in future months, which of the following steps should be taken?
How has ibms stock been doing currently : Select "Investors" (US) on the next page on the right side. Select "Financial Snapshot" on the next page.
What is the value of a put option with strike price : IBM stock currently sells for 84 dollars per share. Over 8 months the price will either go up by 7.5 percent or down by -3.0 percent.

Reviews

Write a Review

Computer Engineering Questions & Answers

  What is the expected time to discover the correct password

Assuming feedback to the adversary flagging an error as each incorrect character is entered, what is the expected time to discover the correct password?

  Define individual project deliverable length

Juan reached the end of his online course program. His family was so proud of him. Juan's wife wanted to throw a party to celebrate Juan's online graduation together with all of his family, and started planning the special event.

  Draw a circuit that retains the logic

Draw a circuit that retains the logic of the accompanying figures, which has a vertical contact (D), so the circuit could be programmed into a PLC.

  Describe a small specific change that you could make

Describe a small specific change that you could make to the code in question 2 that would trigger an error in that phase (and not any previous phase).

  As a member of the information security team at a small

as a member of the information security team at a small college you have been made the project manager to install an

  Explain what each section of code is doing-checking the room

Explain what each section of the code is doing-checking the room, then checking the possible directions in the room.

  Construct the binary search tree

Construct the binary search tree for the following input stream, assuming no balancing or pivoting is done: Frodo, Bilbo, Smaug, Gandalf, Wormtongue, Denethor, Sauron, Galadriel, Aragorn.

  Write a program that creates one pile of marbles

Write a program that creates one pile of marbles with a random number of marbles and decides who starts the game. The program will call userPlay when the user plays and playNovice when it is the computer turns.

  What changes are needed so that a semicolon will be ignored

What changes are needed so that a semicolon will be ignored at the end of the expression but will be an error elsewhere?

  Questionassume this loop is taken many times what is

questionassume this loop is taken many times what is steady-state cpi of this loop on the scalar pipeline discussed in

  Questionsql queriesdownload stovesaccdb access database in

questionsql queriesdownload stoves.accdb access database in doc sharing it has following tables filled with

  List three functions of the amoeba microkernel

List three functions of the Amoeba microkernel. What is it about Amoeba that makes it impossible for clients to tell the difference?

Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd