<?xml version="1.0"?>
<rss version="2.0">
   <channel>
      <title>MLFH Journal Club April 2022 by Moses I</title>
      <link>https://padlet.com/Mojame/n41xn7b0zydb7hph</link>
      <description>Please can each group put up a summary of their thoughts on the assigned paper.</description>
      <language>en-us</language>
      <pubDate>2022-02-16 13:36:59 UTC</pubDate>
      <lastBuildDate>2022-04-29 18:22:26 UTC</lastBuildDate>
      <webMaster>hello@padlet.com</webMaster>
      <image>
         <url>https://padlet.net/icons/png/1f9ee.png</url>
      </image>
      <item>
         <title>Instruction</title>
         <author>Mojame</author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2050859146</link>
         <description><![CDATA[<div>Add more sections by clicking the add button beneath each post per group so its all stacked neatly</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-02-16 13:43:57 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2050859146</guid>
      </item>
      <item>
         <title>Example Group 2</title>
         <author>Mojame</author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2050865284</link>
         <description><![CDATA[<div>You can do a heading and content for each addition</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-02-16 13:46:37 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2050865284</guid>
      </item>
      <item>
         <title></title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2072063737</link>
         <description><![CDATA[]]></description>
         <enclosure url="https://padlet-uploads.storage.googleapis.com/1607164963/bcb4f4b34631f55d2ed53fc4f8ad57dc/JournalClub.docx" />
         <pubDate>2022-03-01 18:29:25 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2072063737</guid>
      </item>
      <item>
         <title>Main Aims</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114796492</link>
         <description><![CDATA[<div>The main aims were to: test the validity of 84 machine learning algorithms to characterize epidemiology of injuries of those Major League Baseball players on the Disabled List using historical data from 2000 to 2017, and predict risk of future injuries. Further aims were to predict the location of the injury as a means to target prevention strategies, and compare the ML algorithms to logistic regression and identify the approach with the greatest predictive capability.</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-26 12:43:16 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114796492</guid>
      </item>
      <item>
         <title>Findings of the study</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114796917</link>
         <description><![CDATA[<div>For predicting future injuries in position players, the top model was ‘top 3 ensemble’ based on highest AUC, giving the best predictions for all types of injuries included in the analysis, except for elbow, which was better modeled with logistic regression. The top three variables for predicting future injury were calculated as: previous injury, weighted cutter runs per 100 pitches, and wins above replacement.<br><br></div><div>For pitchers, the top models for predicting future injuries were random forest and top 3 ensemble, again based on AUC. The latter had a higher degree of accuracy<br><br></div><div>For predicting injuries in four body regions, the top 3 ensemble was the best model for all but the elbow region but the determinants of these injuries could not be calculated in pitchers due to the lower AUCs. This may be a result of limited data.<br><br></div><div>Machine learning models, particularly top 3 ensemble and random forest outperformed logistic regression in 13 of 14 cases.</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-26 12:44:08 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114796917</guid>
      </item>
      <item>
         <title>What are the research questions, and are they relevant for machine learning methods?</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114797741</link>
         <description><![CDATA[<div>The research questions were predictive in nature: data-driven approaches to injury prevention based on historical information, potential location of injury, important variables associated with injuries to players, and the risk of injury. The research also attempted to validate different machine learning algortihms to identify the most accurate.<br><br></div><div>As the questions are focused on predicting future outcomes based on historical data spanning multiple years, many players and sites of injury, machine learning is a relevant method as it offers a robust approach to mining large and diverse datasets to make predictions and identify important indicators.</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-26 12:45:34 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114797741</guid>
      </item>
      <item>
         <title>Which machine learning algorithm was selected? Is it an appropriate method? What are the advantages and disadvantages of this method? You can give reasons.</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114798425</link>
         <description><![CDATA[<div>The algorithm that performed best was the top 3 ensemble, which is a combination of the top three algorithms. A soft voting classifier combines the predictions of different models, taking the probability of each variable and ‘shuffling’ the training data and datapoints which are then passed to each algorithm.<br><br></div><div>The method seems to be appropriate as it combines the best of predictive power from the top-performing models and seems to provide a much higher level of accuracy than individuals models alone. However it may be complicated and mode time-consuming to apply to a dataset, and not entirely straightforward to understand exactly how it works (a practical demonstration would be useful).</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-26 12:46:48 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114798425</guid>
      </item>
      <item>
         <title>Describe the application of each step in the machine learning workflow: data splitting, selection of predictors, model selection, training, optimisation, performance assessment. If these are not available, state so and look for the authors&#39; reasons.</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114799222</link>
         <description><![CDATA[<ul><li>Data splitting: into model testing data – new player with performance and past injury data; and model training data - injury data and performance data<br><br></li><li>Training data are combined and each model is fit to the data<br><br></li><li>The best ML agorithms are determined by calibrating each model against each other and evaluated. The top 3 ensemble approach shuffles the training data and datapoints. Each model calculates the individual prediction with voting aggregator and computes the majority voting for the final prediction (Kumari <em>et al</em>. 2021. An ensemble approach for classification and prediction of diabetes mellitus using soft voting classifier. International Journal of Cognitive Computing in Engineering 2: 40-46). Eventually a final combination of models will be identified.<br><br></li><li>This final model is used on the testing data to predict injury risk and location<br><br></li></ul><div>Models are evaluated by assessing the Area Under the Receiver Operating Charactersitc (ROC) Curve (AUC), % accuracy, F1 score and Brier score loss.</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-26 12:48:17 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114799222</guid>
      </item>
      <item>
         <title>Study limitations</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114799758</link>
         <description><![CDATA[<div>Models for pitchers were not as reliable due to the limited availability of specific data on overuse injuries in modern pitcher databases.<br><br></div><div>Even though the top 3 ensemble approach provided a high level of accuracy and predictability overall, there were some injuries for which a different algorithm, such as random forest, was superior. There is no single algorithm that works for all data and categories of prediction.<br><br></div><div>The granularity of available data was limited, resulting in an inability to predict severity or nature of future injury in a particular body part (e.g. sprain versus tear).<br><br></div><div>It was also not possible to predict the impact of chronic injuries on future injuries.<br><br></div><div>There is also a lack of anatomic specificity to the prediction, which limits the clinical usefulness of the model.<br><br></div><div>The quality, accuracy and longitudinal coverage of the datasets used was limited, either because of uncertainty of reporting (under-reporting), not all time periods were covered, and only one database has been endorsed by Major League Baseball.</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-26 12:49:12 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114799758</guid>
      </item>
      <item>
         <title>Ethical and legal implications of the study</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114800153</link>
         <description><![CDATA[<div>One or more authors declared potential conflicts of interest, having received fees for consulting and/ or other purposes from various companies with interests in the research.<br><br></div><div>Ethical approval was not sought for the study. There is no indication whether the research team were able to identify individuals, took necessary steps to ensure the identities of individuals would be protected from disclosure in any way, nor sought consent from individuals whose data were being used.<br><br></div><div>The US has a sectoral approach to data privacy protection at the federal level, with additional legislation at state-level ergo how personal data are used and protected varies between different states.</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-26 12:49:53 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2114800153</guid>
      </item>
      <item>
         <title>Main aims</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2117061275</link>
         <description><![CDATA[<div>To determine the validity of a machine learning model in predicting the next-season injury risk and anatomic injury location for both position players and pitchers in the Major League Baseball.<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-28 12:46:35 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2117061275</guid>
      </item>
      <item>
         <title>Findings of the study</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2117130870</link>
         <description><![CDATA[<div>For position players, the top 3 ensemble was the best predictive model for future injuries of each anatomic region, with highest AUC 0.76 and the best accuracy at 70%. That except for the elbow, the elbow injuries were best predicted with LR, with an accuracy of 63.0% and an AUC of 0.61.&nbsp;</div><div>&nbsp;</div><div>For pitchers, the models with the highest AUC were random forest and the top 3 ensemble, both with a mean AUC 0.65. The top 3 ensemble model had the highest accuracy at 63.7%.</div><div>&nbsp;</div><div>Back injuries had the highest AUC among both position players and pitchers, at 0.73.&nbsp;<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-28 13:18:28 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2117130870</guid>
      </item>
      <item>
         <title>What are the research questions, and are they relevant for machine learning methods?</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2119000813</link>
         <description><![CDATA[<div>Based on the historical data of MLB players’ injury and performance data between the years of 2000 and 2017, if machine learning can develop an optimised algorithm (better than logistic regression) to predict the next season injury risks and anatomic locations.&nbsp;</div><div>&nbsp;</div><div>Machine learning is a relevant method as it is suitable for operating a big data source with complex data relationships, as well as for generating a predictive model for players’ injury prediction.<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-29 09:41:38 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2119000813</guid>
      </item>
      <item>
         <title>Which machine learning algorithm?</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2119160545</link>
         <description><![CDATA[<div>Six different model algorithms were used: LR, random forest, k-nearest neighbors, Naı¨ve Bayes, XGBoost and top 3 ensemble.</div><div>&nbsp;</div><div>The advanced ML models are superior to LR, as advanced ML models, usually the top 3 ensemble and random forest, outperformed LR in terms of the AUC in 13 of the 14 cases. Specifically, the top 3 ensemble was the model with the highest AUC for predicting next season’s injury risk among position players and pitchers, but random forest was superior in predicting back injuries among pitchers.&nbsp;</div><div>&nbsp;</div><div>With more iterations, the ML algorithm continued to improve.&nbsp; While regression analysis is static and not predictive, especially when more data inputs are added.<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-29 11:45:48 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2119160545</guid>
      </item>
      <item>
         <title>Amy Ogungbemi Journal Club</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2122258433</link>
         <description><![CDATA[<div><strong>1.</strong>&nbsp; &nbsp; &nbsp; <strong>Main aims<br></strong><br></div><div>The aim of this study is to analyse performance data for baseball players and make predictions for future outcomes and prediciting future injuries&nbsp; for players.&nbsp; &nbsp;&nbsp;<br><br></div><div>&nbsp;<br><br></div><div><strong>2.</strong>&nbsp; &nbsp; &nbsp; <strong>Findings of the study.<br></strong><br></div><div>For position players 44% had previous injuries and found that hand and back were most common injuries. And 43.6% of pitchers had previous injuries and most common injuries were shoulder and elbow.&nbsp; The models had accuracy of 0.71-0.8 for position players and 0.61-0.69 for pitchers for future injuries. &nbsp; This reliability is reduced for pitchers due to limited data. &nbsp;<br><br></div><div><strong><em>3.</em></strong>&nbsp; &nbsp; &nbsp; <strong>What are the research questions, and are they relevant for machine learning methods? </strong><strong><em>You can give reasons.<br></em></strong><br></div><div>The research question is predicting injury Logistic regression is used and regression is static and not predictive and is difficult to learn and predict. &nbsp;<br><br></div><div><strong>4.</strong>&nbsp; &nbsp; &nbsp; <strong>Which machine learning algorithm was selected? Is it an appropriate method? You can give reasons. What are the advantages and disadvantages of this method?<br></strong><br></div><div>Different algorithms were used were developed for seven different outputs.&nbsp; A single model is not appropriate for answering all the clinical questions.&nbsp; The models were evaluated for their accuracy.&nbsp; There is different levels of accuracy for different player types (pitcher and positional).&nbsp; For pitcher, random forest and top 3 ensemble had best accuracy and for position players the top 3 ensemble was most accurate. &nbsp;<br><br></div><div>.</div><div><strong><em>5.</em></strong>&nbsp; &nbsp; &nbsp; <strong>Describe the application of each step in the machine learning workflow: data splitting, selection of predictors, model selection, training, optimisation, performance assessment. </strong><strong><em>If these are not available, state so and look for the authors' reasons.&nbsp;<br></em></strong><br></div><div>Separate models were built for position player and pitchers.&nbsp; Different models were build for each clinical outcome.&nbsp; 6 different models were created for each clinical outcome.&nbsp; LR, random forest, k-nearest neighours, naïve bayes, xgboost and top 3 ensemble.&nbsp; The best model was identified and chosen as final model.&nbsp; &nbsp; They used python library and XGBoost. Each model used a 10 k-fold.&nbsp; 90 percent of data was used to train the model and 10 percent used to test it.&nbsp; This step is repeated 10 times. &nbsp;<br><br></div><div><strong>6.</strong>&nbsp; &nbsp; &nbsp; <strong>Study limitations<br></strong><br></div><div>One limitation is the granularity of available data.&nbsp; That is we were unable to analyse the level of detail.&nbsp; For this it is not possible to ascertain the level of injury (sprain vs complete tear of a ligament).&nbsp; There is limited anatomic specificity which limits how clinically useful it will be.&nbsp; The data was unable to show the impact of chronic conditions.&nbsp; The data used is a limitation as owned by private companies and 3 databases were not regulated.&nbsp; Some data was not as specific as other data.<br><br></div><div><strong>7.</strong>&nbsp; &nbsp; &nbsp; <strong>Ethical and legal implications of the study<br></strong><br></div><div>The ethics of predicting an injury can affect a players potential worth as it can create a quantifiable risk of injury to a player. &nbsp;<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-30 20:26:18 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2122258433</guid>
      </item>
      <item>
         <title>Workflow</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2123316099</link>
         <description><![CDATA[<div>Data splitting: 90% of the data be used to train the model, and the remaining 10% is used to test the model in an unbiased fashion. </div><div>&nbsp;</div><div>Selection of predictors: age, performance data, injury history, and DL data from 17 seasons were predictive of next-season injuries.</div><div><br></div><div>Model selection: each model is evaluated by Accuracy, AUC, F1 score, Brier Score Loss, the highest AUC primarily determines the best performing model.</div><div>&nbsp;</div><div>Training: Each model uses 10 k-folds to train 90% of the data and the remaining 10% is used to test the model. This step is repeated 10 times.</div><div>&nbsp;</div><div>Optimisation: it is not available in this article.<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-31 10:32:13 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2123316099</guid>
      </item>
      <item>
         <title>Study limitations</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2123330129</link>
         <description><![CDATA[<div>The granularity of available data is limited.</div><div>Lack of anatomic specificity of the data prediction algorithm that limited immediate clinical utility of such a model.</div><div>It is not possible to capture the impact that chronic, lingering injuries may have on future injuries.<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-03-31 10:43:36 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2123330129</guid>
      </item>
      <item>
         <title>Journal Club</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125067116</link>
         <description><![CDATA[<div><strong>1. Main Aims</strong><br><br></div><div>To determine the validity of an ML model in predicting next seasons injury risk and anatomic injury location for both position players and pitchers in the MLB.<br><br></div><div><strong>2. Findings of the study</strong><br><br></div><div>Advanced ML models generally outperformed logistic regression and demonstrated fair capability in predicting publicly reportable next season injuries, including the anatomic region for position players, although not for pitchers.&nbsp;<br><br></div><div>Advanced ML models outperformed logistic regression in 13 of 14 cases.<br><br></div><div><strong>3. Research questions (relevant for machine learning?)</strong><br><br></div><div>Can we characterise the epidemiology of injury trends on the disabled list (DL) from 2000 to 2017?<br><br></div><div>Yes suitable, unsupervised ML model.<br><br></div><div>Can we determine the validity of an ML model in predicting the injury risk for the subsequent year and anatomic injury location?<br><br></div><div>Yes suitable, supervised ML model.<br><br></div><div>Can we compare the performance of modern ML algorithms versus logistic regression analyses?<br><br></div><div>Not suitable for ML, just a comparison of performance.<br><br></div><div><strong>4. Which machine learning algorithm was selected?</strong><br><br></div><div>Random forest was selected and top 3 ensemble (random forest, logistic regression &amp; XGBoost), but others were tested.<br><br></div><div>a. Is it an appropriate method?<br><br></div><div>Note sure, it seems that different models or combinations of models had different accuracy for different parts of the analysis. The random forest model worked best for back injuries amongst pitchers. The top 3 ensemble had the best area under the curve (AUC) results for predicting next seasons injury risk amongst position players and pitchers.<br><br></div><div>b. What are the advantages and disadvantages?<br><br></div><div>Random forest - Black box – don’t know exactly how the outcome was calculated. &nbsp;<br><br></div><div>Top 3 ensemble – potentially a very complicated ML model again lacking interpretability&nbsp;<br><br></div><div><strong>5. Describe the application of each step in the machine learning model.</strong><br><br></div><div>a. Data Splitting<br><br></div><div>The data was split 90% training / 10% test<br><br></div><div>b. Selection of Predictors<br><br></div><div>Age, performance data, professional injury history and DL list were categories within the data set. The importance of the variables were assessed and injury was found to be the most important.<br><br></div><div>c. Model Selection<br><br></div><div>Multiple models were chosen for assessment.<br><br></div><div>d. Training<br><br></div><div>The models were trained and evaluated using cross validation 10 fold.<br><br></div><div>e. Optimisation<br><br></div><div>Feature importance was calculated using the Gini importance metric.<br><br></div><div>f. Performance assessment<br><br></div><div>Area under the curve (AUC) was used to choose the machine learning model with the best accuracy.<br><br></div><div><strong>6 Study limitations</strong><br><br></div><div>Data bases are updated from multiple sources and could contain multiple errors. Databases were used that were not regulated by MLB. Couldn’t be specific about different injuries ie. Elbow sprain or complete tear<br><br></div><div>&nbsp;<br><br></div><div><strong>7. Ethical and legal implications of the study</strong><br><br></div><div>Data used without accuracy. Could affect the players earnings and contracts.&nbsp;<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-04-01 08:27:50 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125067116</guid>
      </item>
      <item>
         <title>Journal Club</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125167415</link>
         <description><![CDATA[<div>1.&nbsp; &nbsp; &nbsp; Main aims<br><br></div><div>Characterize the epidemiology of injury trends of an ML model in predicting the injury risk for the subsequent year and anatomic injury location.<br><br></div><div>Compare the performance of modern ML algorithms versus LR analyses. &nbsp;<br><br></div><div>2.&nbsp; &nbsp; &nbsp; Findings of the study.<br><br></div><div>The ML models outperformed logistic regression in 13 of 14 cases. The back injuries had the highest AUC among players. The AUC for predicting next-season injuries was 0.76 among position players and 0.65 among pitchers.&nbsp;<br><br><br></div><div><em>3.</em>&nbsp; &nbsp; &nbsp; What are the research questions, and are they relevant for machine learning methods? <em>You can give reasons.<br></em><br></div><div>Can the charactierization of the epidemiology of injury trends of MBL players be used in a ML model for predcting the injury risk for the next year and the anatomic injury location?<br><br></div><div>Can this ML model outperform the linear regression analyses?<br><br></div><div>These are relevant questions that try to find a predictive model, an objective of machine learning. The methodology also includes a vast number of data points, a characteristic that makes machinle learning models specially useful.&nbsp;<br><br><br></div><div>4.&nbsp; &nbsp; &nbsp; Which machine learning algorithm was selected? Is it an appropriate method? You can give reasons. What are the advantages and disadvantages of this method?<br><br></div><div>&nbsp;Top 3 ensamble for position players.&nbsp;<br><br></div><div>Top 3 ensemble and Random Forest for the pitchers.<br><br></div><div>Ensemble methods create different models and combine them to obtain improved results. The models are selected by “voting” (as is the case in this paper) or by “averaging”. While they can improve the predictive values , the interpretability also diminishes as it is a “black box”.<em>&nbsp;<br></em><br></div><div>6.&nbsp; &nbsp; &nbsp; Study limitations<br><br></div><div>The data did not contained nuanced injury characteristics for a more accurate diagnosis.<br><br></div><div>Team-reported injuries are generally acute and severe, so the impact of chronic, lingering injuries could not be evaluated.&nbsp;<br><br></div><div>Lack of anatomic specificity on the prediction algorithm.&nbsp;<br><br></div><div>The databases that are privately owned&nbsp; are not regulated by MLB. The one endorsed by the MLB started in 2015.<br><br></div><div>Less predictive value for the pitcher position compared with the other players, as the pitcher data is less specific in terms of predictive variables.&nbsp;<br><br></div><div><br></div><div>7.&nbsp; &nbsp; &nbsp; Ethical and legal implications of the study<br><br></div><div>It can affect the future employability and value of players identified as “at risk” of injury. Accessing more specific data could require the approval of the MLB Players Association.&nbsp;<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-04-01 09:48:35 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125167415</guid>
      </item>
      <item>
         <title>Main Aims</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125186074</link>
         <description><![CDATA[<div>The study was aimed at leveraging available analytics to permit data-driven injury prevention strategies and informed decisions with the following objectives in mind:<br>&nbsp; &nbsp;- Characterize the epidemiology of injury trends on the DL from 2000 to 2017<br>&nbsp; &nbsp;- Determine the validity of an ML model in predicting the injury risk for the subsequent year and anatomic injury location.</div><div>&nbsp; &nbsp;- Compare the performance of modern ML algorithms versus LR analyses.<br><br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-04-01 10:07:20 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125186074</guid>
      </item>
      <item>
         <title>Study limitations:.</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125224243</link>
         <description><![CDATA[<div>Below are the Limitation of this study:<br>&nbsp; 1.&nbsp; The granularity of available data. The database lacks prior injury data&nbsp; and was provided&nbsp; without the context of performance metrics.<br>  2.&nbsp; Inability to capture the impact that chronic, lingering injuries may have on future injuries, as team-reported injuries are generally acute and severe enough to withdraw players from games.</div><div>  3.&nbsp; &nbsp;Lack of anatomic specificity of the data prediction algorithm used.</div><div>  4.&nbsp; &nbsp;Sources of input of the databases used in  obtaining the MLB player injury history and performance data is either prone to inaccuracies or not regulated <br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-04-01 10:45:53 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125224243</guid>
      </item>
      <item>
         <title>Machine Learning algorithm</title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125251190</link>
         <description><![CDATA[<div>Logistic Regression&nbsp;<br>Random Forest<br>K-nearest neighbors<br>Naive Bayes<br>XGBoost&nbsp;<br>Top 3 ensemble&nbsp;</div>]]></description>
         <enclosure url="" />
         <pubDate>2022-04-01 11:14:52 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125251190</guid>
      </item>
      <item>
         <title></title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125370618</link>
         <description><![CDATA[]]></description>
         <enclosure url="https://padlet-uploads.storage.googleapis.com/1609015177/9520dcd4009e008dc556d8c49a9bec5a/Journal_club.docx" />
         <pubDate>2022-04-01 12:53:06 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2125370618</guid>
      </item>
      <item>
         <title></title>
         <author></author>
         <link>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2143923527</link>
         <description><![CDATA[<div>1.&nbsp; &nbsp; &nbsp; &nbsp;Main aims<br><br></div><div>The main aims of this study were to apply machine learning methodologies to predict the likelihood of injuries in baseball players. Specifically, the authors aimed to develop a machine learning model to predict the risk and site of injury, and compare this to more established analytical models (namely logistic regression analysis).&nbsp;<br><br></div><div>&nbsp;<br><br></div><div>2.&nbsp; &nbsp; &nbsp; &nbsp;Findings of the study.<br><br></div><div>In this study, multiple machine learning models were developed to predict distinct risks (e.g., risk of overall injury, risk of knee/back/ankle/etc. injury), including a combination of the best performing models. Looking at these models, the best predictor of future injury was previous injury. Generally, the combined models had the best accuracy as measured by area under the curve, Brier score loss, and F1 score. Compared to logistic regression analysis, machine learning models were generally better at predicting future injury (13/14 models out-performing LR); this was especially true when combining the best-performing models.<br><br></div><div>&nbsp;<br><br></div><div><em>3.</em>&nbsp; &nbsp; &nbsp; &nbsp;What are the research questions, and are they relevant for machine learning methods? <em>You can give reasons.<br></em><br></div><div>The research question in this study was to assess the applicability of machine learning methodologies to real-world baseball data to predict injury. This is relevant for machine learning methods. Developing machine learning models using real-world data to predict injury highlights the real-world application of machine learning methods; by focusing on injuries necessitating time off play, these models could theoretically lead to significant cost savings as well. The nature of baseball data (thanks to sabermetrics) meant a large pool of real-world data was available across a number of years, and this would be a suitable dataset on which to trial machine learning methods.<br><br></div><div>&nbsp;<br><br></div><div>4.&nbsp; &nbsp; &nbsp; &nbsp;Which machine learning algorithm was selected? Is it an appropriate method? You can give reasons. What are the advantages and disadvantages of this method?<br><br></div><div>The combination models (“top 3 ensemble”) using the best-performing machine learning models were selected for the final analysis. This was appropriate in terms of developing a highly accurate predictive model, as this had consistently good accuracy in predicting overall and site-specific injury compared to other models.<br><br></div><div>One main advantage of this method was its high accuracy while remaining similar in performance to other models. One disadvantage of this was the relative lack of transparency compared to other models, which could limit its reproducibility.<br><br></div><div>&nbsp;<br><br></div><div><em>5.</em>&nbsp; &nbsp; &nbsp; &nbsp;Describe the application of each step in the machine learning workflow: data splitting, selection of predictors, model selection, training, optimisation, performance assessment. <em>If these are not available, state so and look for the authors' reasons.&nbsp;<br></em><br></div><div>The data were split into distinct 10% sections, with 9 of these used to train the model and the final 1 used to test the model. No predictors were specifically selected; instead, all available data were added to the model. The models were trained using a separate 10% of the data as outlined above using a 10 k-fold strategy. These models were calibrated by assessing their performance against one another, but were not clearly optimised. The performance of the models was assessed using accuracy, area under the ROC curve, F1 score, and Brier score loss.&nbsp;<br><br></div><div><em>&nbsp;<br></em><br></div><div>6.&nbsp; &nbsp; &nbsp; &nbsp;Study limitations<br><br></div><div>This study had a few limitations. The “black box” of the top 3 ensemble model, for example, limits the reproducibility of the study’s findings and may similarly limit its applicability to other situations.&nbsp; As well, as noted by the study authors, a single model may not be the best way to address the clinical questions – as seen in the study findings, there was variation in the ability of the models to predict specific anatomical sites. This may limit its applicability in a real-world setting, as not all injuries are similarly debilitating and may necessitate distinct interventions.<br><br></div><div>There are also the usual limitations which come with using big data sources, such as the ethical considerations of the data, the accuracy and reliability of the data, and the possibility of missing important qualitative information. Specifically relevant to this study, the lack of detailed information about the injuries within the data may have affected the interpretability of the findings.&nbsp;<br><br></div><div>&nbsp;<br><br></div><div>7.&nbsp; &nbsp; &nbsp; &nbsp;Ethical and legal implications of the study<br><br></div><div>There are ethical considerations to consider when using data from individuals. Although the information is publicly available, consent was not explicitly taken from the players for their data to be used to form the models used in this study. The models used may also indirectly affect the future earnings of players based on their risk of future injury.<br><br></div>]]></description>
         <enclosure url="" />
         <pubDate>2022-04-14 14:10:14 UTC</pubDate>
         <guid>https://padlet.com/Mojame/n41xn7b0zydb7hph/wish/2143923527</guid>
      </item>
   </channel>
</rss>
