<?xml version="1.0"?>
<rss version="2.0">
   <channel>
      <title>HW3 Discussion and Q&amp;A by </title>
      <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv</link>
      <description>Please ask any HW3 questions you have here. If we receive a HW3-related question in writing, we&#39;d share the question and answer here (anonymously).</description>
      <language>en-us</language>
      <pubDate>2023-04-04 14:00:22 UTC</pubDate>
      <lastBuildDate>2023-04-29 10:31:53 UTC</lastBuildDate>
      <webMaster>hello@padlet.com</webMaster>
      <image>
         <url></url>
      </image>
      <item>
         <title>Question about &quot;current baseline/practice&quot; </title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2543798708</link>
         <description><![CDATA[<div><strong>Question: </strong>It says in the assignment that Company X is currently not predicting default, and we're then confused on the meaning behind current baseline if they're not actually predicting anything as "current practice".</div><div>Is the baseline current practice just the simplest model where we get a value that we can compare our model with? Or does it mean that we use Company X's current practice? In HW1 the current practice was given, but this time it sounds like they just don’t have a current practice, so we're a bit confused. <br><br><strong>Answer: </strong>An example mentioned in class is that someone walks into the doctor's office and that doctor tests everyone for breast cancer. I emphasized how that is in itself making a prediction. The same thing applies here.&nbsp;</div><div>To answer your two questions with a question mark: It means that you use the Company X's current practice.&nbsp;</div><div>In our (last) review class, it's mentioned that often the current baseline has to be inferred and figured out and is not explicitly stated as in HW1. In real life, that's almost always the case.</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-04 14:59:14 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2543798708</guid>
      </item>
      <item>
         <title>Question 1 states that suspending spending privileges can be an action/intervention, then in question 3 after putting up the profit matrix asks &quot; Are there qualitative elements you cannot measure? If so, mention them, and explain how this will affect your eventual model valuation&quot;. We then thought to mention the administrative efforts/cost, because measures like customer satisfaction would be difficult to quantify without experts. Should we look away from this?</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2544886953</link>
         <description><![CDATA[<div><strong>Answer</strong>: the question states "there is no cost/benefit associated with customer goodwill" before it specifies question 1. That means assume that. So you can disregard customer goodwill/dissatisfaction costs completely. Question 3.ii then asks you to specify qualitative elements (here, you shouldn't mention customer goodwill, because the earlier assumption asked you to disregard it; <em>Sidepoint: HOWEVER, administrative costs as a qualitative element is a fair answer, as can be independent of customer goodwill; e.g. you may say that if the customer wasn't to default but we suspended their spending privileges, we'd have to send a letter and apologize to them.</em>), and then 3.iii(1) asks for the numeric values in each profit matrix cell, and asks if there is one you cannot measure, then mention it and explain how it will affect your model evaluation. Here you can mention the administrative costs that I mentioned above, and specify how it will affect your model. Just don't forget to specify the numeric values in each profit matrix cell too.&nbsp;</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-05 11:21:18 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2544886953</guid>
      </item>
      <item>
         <title>If &quot; there is no cost/benefit associated with customer goodwill&quot; or customer dissatisfaction, how can we in question 1 formulate action/interventions.</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2544889114</link>
         <description><![CDATA[<div><strong>Answer</strong>: Actions don't rely on you having customer goodwill; An action that the assignment suggests is "suspending spending privileges", that is the action. If customer goodwill existed (which we've asked you to assume it doesn't), it would have contributed to qualitative elements in the matrix, it wouldn't have been an action itself.&nbsp;<br>The assignment only expects you to value "suspending spending privileges", but to mention 3 other potential actions and justify your choice (specifying their limitations would be a plus here).&nbsp;</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-05 11:24:28 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2544889114</guid>
      </item>
      <item>
         <title>Hello! We have several questions for this amazing project :D</title>
         <author></author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2546261779</link>
         <description><![CDATA[<div>1) Is the balance in the ‘bill-amt’ column, the ending balance or the beginning balance?<br>2) For Q3(iii), do we need to get the final value for the answer or can we just put the generic equation, such as 2% of carried balance instead of the final value.&nbsp;<br>3) For Q4, to prepare the data for continuous prediction, do we need to add new attributes to the dataset? For example, how much is the actual default for each customer.<br>4) For Q4, is it possible to consider partial default as ‘default’? Or it is always 100% default. For example, the outstanding balance of customer A is 10,000 NOK and next month he paid only 6,000 NOK. Do we call 4,000 NOK default?</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-06 13:17:32 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2546261779</guid>
      </item>
      <item>
         <title>Questions about confusion matrix basline model</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2546531178</link>
         <description><![CDATA[<div><strong>Question 1</strong>: We are to create a confusion matrix basline model , isnt it ?</div><div><strong>Answer 1</strong>: Yes, that’s right.</div><div><strong>Question 2</strong>: Is it the baseline model in a formula sort of way like you mentioned in lecture 11. &nbsp;</div><div><strong>Answer 2</strong>: The point of the formula was different. it was to say that sometimes a profit matrix is to be specified as a constant value, or as depending on other parameters. it's for you to figure out which applies for the baseline in this case.<br> <strong>Question 3</strong>: If it is the confusion matrix method , Should we :<br>1.logically assume values that go into the confusion matrix ? or</div><div>&nbsp;2.should we go with some sort of numbers from within the assignment ?&nbsp;</div><div><strong>Answer 3</strong>: Use values from the assignment, and where you need to make an assumption, make it and state it.</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-06 18:15:28 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2546531178</guid>
      </item>
      <item>
         <title>We are a bit confused by the dataset. Is the bill columns accumulated or monthly? </title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2548570792</link>
         <description><![CDATA[<div><strong>Answer</strong>: The bill columns are accumulated. For example, if BILL_AMT4 JUNE&nbsp; is 3272 and BILL_AMT3 JULY is 2682, then 2682 NT dollars (but not 3272 + 2682 NT dollars) is the total amount owed to the bank before August. So, 2682 is the accumulated bill by the end of July.&nbsp;</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-10 07:09:38 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2548570792</guid>
      </item>
      <item>
         <title>Question about valuation</title>
         <author></author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2549986975</link>
         <description><![CDATA[<div>Hi,<br>In the last question we have to do a valuation of the model. I've watched lecture 12 but that dataset is very different from what we have. Is there any other resources that would be helpful when doing the valuation?</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-11 10:47:56 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2549986975</guid>
      </item>
      <item>
         <title>Question about Q4</title>
         <author></author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2550687904</link>
         <description><![CDATA[<div>Hi,<br>Is next month's default amount meaning in Q4 is the predicted BILL Amount on October*prediction value from classification model?</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-11 20:29:32 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2550687904</guid>
      </item>
      <item>
         <title>A follow-up question from students about Q4</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2551299075</link>
         <description><![CDATA[<div><strong>Question</strong>: How to interpret the negative values? Are the negative values in the dataset people who have overpaid their account or is negative when someone else pays that persons balance. And if it is someone paying for them how do we know their actual amount owed?<br><strong>Answer</strong>: The negative value means cardholders overpaid their account. 5000 and -2000 are the balances of two different cardholders. The hint asks whether predicting negative values (e.g. -2000) would make any sense. So you can think about this hint when answering Q4. &nbsp;</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-12 07:43:52 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2551299075</guid>
      </item>
      <item>
         <title>Question about data inconsistency</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2551306327</link>
         <description><![CDATA[<div><strong>Question</strong>: Hello, we are currently looking at the dataset for the ML exam. We have noticed several enological values in the dataset. for example there are several customers with 0 bill AMT and Pay AMT but still labeled as default next month, we also have customers with the same values which are not labeled as defaults next month. Further the Pay column seem inconsistent where there are several -2 values that should be -1 and opposite. We are thinking of how to deal with this, either we try and adjust the dataset to remove most of these inconsistencies or should we take the dataset as it is and just include a statement in our task discussing the inconsistencies in the data and how that can affect the model.<br><strong>Answer</strong>: You are right, the dataset is imperfect and there are some data inconsistencies. I guess you are asking question related to Q6. So you should specify the data issues you notice, state any recoding/preprocessing you think would be useful to deal with the data, and think whether deriving new features could be useful.</div><div>The question asks how you would preprocess these data, to extract more meaningful information/signal from the dataset. So the key point is how you would deal with the data issue and make it make sense.</div><div>Adjusting the dataset is one way to do it, you can also deal with it in other ways as long as it makes sense.</div><div>It would be better if you can present how you preprocess the data instead of how the data issues can affect the model.</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-12 07:51:24 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2551306327</guid>
      </item>
      <item>
         <title>For question 4 is it hypothetical or should we actually create a model? So then we would have two models one for classification and one for regression predictions and our baseline?</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2552860018</link>
         <description><![CDATA[<div>No, you don’t have to create a regression model in Q4. The question is “How would you prepare the data for use in a regression model?”, so you only need to clarify how you prepare the data for the regression model with the raw data.&nbsp;</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-13 09:25:43 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2552860018</guid>
      </item>
      <item>
         <title>Question about downsampling</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2552865197</link>
         <description><![CDATA[<div><strong>Question</strong>: When there is imbalance and we decide to downsample for instance. Do we make the majority class equal to the minority class as you had shown in class/slides? Also, then why did you only downsample only 10% of the data in the DR example, is the dataset still not imbalanced?&nbsp;</div><div>Also, In thes slides, basically you have upsampled the minority class and downsampled the majority class to make it equal (500/500). Should we also do both? or is doing one enough?&nbsp;</div><div>Also, what is the safe ratio that we are looking for when we do any of these? <br><br><strong>Answer</strong>: As you mentioned, we can upsample the minority class or downsample the majority class to make it equal (500/500), so you can choose one of them and don’t need to do both.</div><div>I didn’t fully understand the question “why did you only downsample only 10% of the data in the DR example?” and “safe ratio”. I presume you mean the downsampling we did in DataRobot, where we resampled an extremely imbalanced European Credit Card dataset to 1:9 (10%). Here the 10% means the percentage of the minority, and we should resample the training data to make it 5:5.&nbsp;</div><div>With a train-valuation-holdout split, you should resample your training data and get resampled data, so the percentage of the resampling data is the percentage of your training data.</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-13 09:32:33 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2552865197</guid>
      </item>
      <item>
         <title>Question about Model Valuation</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2552870486</link>
         <description><![CDATA[<div><strong>Question 1:&nbsp;</strong>I have a question regarding HW3. In the 8<sup>th</sup> question, it asks the ‘total expected value of the action and model’. Are we supposed to get the total profit through the profit matrix or just an explanation would be enough? (Because we can’t get the value in FN cell as we don’t know the value of new transactions to calculate the ‘fee on new charges’)&nbsp;</div><div>Also, if we are resampling the data, as per Lecture 10 (Page 20), we should evaluate the model on the initial train data set. I have a doubt that when we should use the test data set. Can you please clarify that as well?<br><br><strong>Answer</strong> <strong>1</strong>: Yes, you are supposed to get the total expected value through the profit matrix.&nbsp;</div><div>The value of ‘new charges’ is not provided in the raw data, so your task is to make simplifying assumptions to assume values for these missing variables based on some domain knowledge of bank customer spending. Given all the previous bill statements for the person that we see, what would likely be missing values in the FN cell? Making reasonable assumptions is part of data science (machine learning for business), if you could do so, explain it, and measure the elements in the profit matrix, then I think you can find the way to calculate the ‘total expected value of the action and model’ (see lecture 12 videos). You're graded on coming up with a reasonable approximation. There's not just one solution, so choose the best approximation you can think of.</div><div>In HW3, you should be able to combine resampling and use the train-valuation-holdout (Valuing Action on a 3rd fold) approach we taught. The best model trained on the resampled data would be your Model* (see lecture 12 videos). Then, you VALUE your model* on the valuation fold to find the best action* (e.g. threshold; watch lecture 12) with given model*. Next, you take the best model* and action* and report the value on the holdout set. The test data set is the holdout set, and it will be used in the end to evaluate the performance of Model* with action*.<br><br><strong>Question 2</strong>: Q3.a.iii.1. says "Are there qualitative elements you cannot measure? If so, mention them, and explain how this will affect your eventual model valuation."</div><div>Since it says "eventual model valuation", aren't we supposed to get the confusion matrix and profit matrix for the holdout set (after valuation) in earlier questions 3.a.i &amp; ii?<br><br><strong>Answer 2</strong>: In Q3, you can show a confusion matrix for any of the full dataset, valuation set or holdout set.</div><div>Also, Q3 asks you to specify the "qualitative elements in each cell of the profit matrix" (which can be static), and then specify how you would "calculate numeric values in each profit matrix cell" (the formula would be static). <strong>Throughout lecture 11</strong>, it’s mentioned that some profit matrices would consist of numbers, some would consist of an equation to calculate profit cell matrix values. In Q3 you are able to show the equation of each cell, and then describe how you calculate it.&nbsp;</div><div>So you can have the confusion matrix and profit matrix for any of the full dataset, valuation set or holdout set, and you can show one of the three datasets.&nbsp;</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-13 09:39:41 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2552870486</guid>
      </item>
      <item>
         <title>Confusion matrix 3i.</title>
         <author></author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2555658515</link>
         <description><![CDATA[<div>Should we use confusion matrix for "suspend spending privileges" in 3a or baseline model? Since it asks "For the action in (1.a), specify the following valuation models" in the question we think we should use c.matrix for the action (suspend spending privileges). Is that right?</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-16 09:09:38 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2555658515</guid>
      </item>
      <item>
         <title>I want to check if I&#39;m thinking correctly:  If the customer spends 6000 on new transaction, did they actually buy something for 5 820 and they pay 180 on transactions?</title>
         <author>meichen8_1</author>
         <link>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2556525433</link>
         <description><![CDATA[<div>If the customer spends 6000 on new transaction, then they actually buy something for 6000. The transaction fee 180 will be paid by the merchant.</div>]]></description>
         <enclosure url="" />
         <pubDate>2023-04-17 07:04:10 UTC</pubDate>
         <guid>https://padlet.com/meichen8_1/tttfcixn9023wbgv/wish/2556525433</guid>
      </item>
   </channel>
</rss>
