LibCoder

Learning Data Science: Data Wrangling, Exploration, Visualization, and Modeling with Python

1C Agda Big Data/DataScience
Learning Data Science: Data Wrangling, Exploration, Visualization, and Modeling with Python
Дата выхода: 2023
Издательство: O’Reilly Media, Inc.
Количество страниц: 597
Размер файла: 5,6 МБ
Тип файла: PDF
Добавил: LibCoder
Оглавление
CoverCopyrightTable of ContentsPrefaceExpected Background KnowledgeOrganization of the BookConventions Used in This BookUsing Code ExamplesO’Reilly Online LearningHow to Contact UsAcknowledgmentsPart I. The Data Science LifecycleChapter 1. The Data Science LifecycleThe Stages of the LifecycleExamples of the LifecycleSummaryChapter 2. Questions and Data ScopeBig Data and New OpportunitiesExample: Google Flu TrendsTarget Population, Access Frame, and SampleExample: What Makes Members of an Online Community Active?Example: Who Will Win the Election?Example: How Do Environmental Hazards Relate to an Individual’s Health?Instruments and ProtocolsMeasuring Natural PhenomenaExample: What Is the Level of CO2 in the Air?AccuracyTypes of BiasTypes of VariationSummaryChapter 3. Simulation and Data DesignThe Urn ModelSampling DesignsSampling Distribution of a StatisticSimulating the Sampling DistributionSimulation with the Hypergeometric DistributionExample: Simulating Election Poll Bias and VarianceThe Pennsylvania Urn ModelAn Urn Model with BiasConducting Larger PollsExample: Simulating a Randomized Trial for a VaccineScopeThe Urn Model for Random AssignmentExample: Measuring Air QualitySummaryChapter 4. Modeling with Summary StatisticsThe Constant ModelMinimizing LossMean Absolute ErrorMean Squared ErrorChoosing Loss FunctionsSummaryChapter 5. Case Study: Why Is My Bus Always Late?Question and ScopeData WranglingExploring Bus TimesModeling Wait TimesSummaryPart II. Rectangular DataChapter 6. Working with Dataframes Using pandasSubsettingData Scope and QuestionDataframes and IndicesSlicingFiltering RowsExample: How Recently Has Luna Become a Popular Name?AggregatingBasic Group-AggregateGrouping on Multiple ColumnsCustom Aggregation FunctionsPivotingJoiningInner JoinsLeft, Right, and Outer JoinsExample: Popularity of NYT Name CategoriesTransformingApplyExample: Popularity of “L” NamesThe Price of ApplyHow Are Dataframes Different from Other Data Representations?Dataframes and SpreadsheetsDataframes and MatricesDataframes and RelationsSummaryChapter 7. Working with Relations Using SQLSubsettingSQL Basics: SELECT and FROMWhat’s a Relation?SlicingFiltering RowsExample: How Recently Has Luna Become a Popular Name?AggregatingBasic Group-Aggregate Using GROUP BYGrouping on Multiple ColumnsOther Aggregation FunctionsJoiningInner JoinsLeft and Right JoinsExample: Popularity of NYT Name CategoriesTransforming and Common Table ExpressionsSQL FunctionsMultistep Queries Using a WITH ClauseExample: Popularity of “L” NamesSummaryPart III. Understanding The DataChapter 8. Wrangling FilesData Source ExamplesDrug Abuse Warning Network (DAWN) SurveySan Francisco Restaurant Food SafetyFile FormatsDelimited FormatFixed-Width FormatHierarchical FormatsLoosely Formatted TextFile EncodingFile SizeThe Shell and Command-Line ToolsTable Shape and GranularityGranularity of Restaurant Inspections and ViolationsDAWN Survey Shape and GranularitySummaryChapter 9. Wrangling DataframesExample: Wrangling CO2 Measurements from the Mauna Loa ObservatoryQuality ChecksAddressing Missing DataReshaping the Data TableQuality ChecksQuality Based on ScopeQuality of Measurements and Recorded ValuesQuality Across Related FeaturesQuality for AnalysisFixing the Data or NotMissing Values and RecordsTransformations and TimestampsTransforming TimestampsPiping for TransformationsModifying StructureExample: Wrangling Restaurant Safety ViolationsNarrowing the FocusAggregating ViolationsExtracting Information from Violation DescriptionsSummaryChapter 10. Exploratory Data AnalysisFeature TypesExample: Dog BreedsTransforming Qualitative FeaturesThe Importance of Feature TypesWhat to Look For in a DistributionWhat to Look For in a RelationshipTwo Quantitative FeaturesOne Qualitative and One Quantitative VariableTwo Qualitative FeaturesComparisons in Multivariate SettingsGuidelines for ExplorationExample: Sale Prices for HousesUnderstanding PriceWhat Next?Examining Other FeaturesDelving Deeper into RelationshipsFixing LocationEDA DiscoveriesSummaryChapter 11. Data VisualizationChoosing Scale to Reveal StructureFilling the Data RegionIncluding ZeroRevealing Shape Through TransformationsBanking to Decipher RelationshipsRevealing Relationships Through StraighteningSmoothing and Aggregating DataSmoothing Techniques to Uncover ShapeSmoothing Techniques to Uncover Relationships and TrendsSmoothing Techniques Need TuningReducing Distributions to QuantilesWhen Not to SmoothFacilitating Meaningful ComparisonsEmphasize the Important DifferenceOrdering GroupsAvoid StackingSelecting a Color PaletteGuidelines for Comparisons in PlotsIncorporating the Data DesignData Collected Over TimeObservational StudiesUnequal SamplingGeographic DataAdding ContextExample: 100m Sprint TimesCreating Plots Using plotlyFigure and Trace ObjectsModifying LayoutPlotting FunctionsAnnotationsOther Tools for VisualizationmatplotlibGrammar of GraphicsSummaryChapter 12. Case Study: How Accurate Are Air Quality Measurements?Question, Design, and ScopeFinding Collocated SensorsWrangling the List of AQS SitesWrangling the List of PurpleAir SitesMatching AQS and PurpleAir SensorsWrangling and Cleaning AQS Sensor DataChecking GranularityRemoving Unneeded ColumnsChecking the Validity of DatesChecking the Quality of PM2.5 MeasurementsWrangling PurpleAir Sensor DataChecking the GranularityHandling Missing ValuesExploring PurpleAir and AQS MeasurementsCreating a Model to Correct PurpleAir MeasurementsSummaryPart IV. Other Data SourcesChapter 13. Working with TextExamples of Text and TasksConvert Text into a Standard FormatExtract a Piece of Text to Create a FeatureTransform Text into FeaturesText AnalysisString ManipulationConverting Text to a Standard Format with Python String MethodsString Methods in pandasSplitting Strings to Extract Pieces of TextRegular ExpressionsConcatenation of LiteralsQuantifiersAlternation and Grouping to Create FeaturesReference TablesText AnalysisSummaryChapter 14. Data ExchangeNetCDF DataJSON DataHTTPRESTXML, HTML, and XPathExample: Scraping Race Times from WikipediaXPathExample: Accessing Exchange Rates from the ECBSummaryPart V. Linear ModelingChapter 15. Linear ModelsSimple Linear ModelExample: A Simple Linear Model for Air QualityInterpreting Linear ModelsAssessing the FitFitting the Simple Linear ModelMultiple Linear ModelFitting the Multiple Linear ModelExample: Where Is the Land of Opportunity?Explaining Upward Mobility Using Commute TimeRelating Upward Mobility Using Multiple VariablesFeature Engineering for Numeric MeasurementsFeature Engineering for Categorical MeasurementsSummaryChapter 16. Model SelectionOverfittingExample: Energy ConsumptionTrain-Test SplitCross-ValidationRegularizationModel Bias and VarianceSummaryChapter 17. Theory for Inference and PredictionDistributions: Population, Empirical, SamplingBasics of Hypothesis TestingExample: A Rank Test to Compare Productivity of Wikipedia ContributorsExample: A Test of Proportions for Vaccine EfficacyBootstrapping for InferenceBasics of Confidence IntervalsBasics of Prediction IntervalsExample: Predicting Bus LatenessExample: Predicting Crab SizeExample: Predicting the Incremental Growth of a CrabProbability for Inference and PredictionFormalizing the Theory for Average Rank StatisticsGeneral Properties of Random VariablesProbability Behind Testing and IntervalsProbability Behind Model SelectionSummaryChapter 18. Case Study: How to Weigh a DonkeyDonkey Study Question and ScopeWrangling and TransformingExploringModeling a Donkey’s WeightA Loss Function for Prescribing AnestheticsFitting a Simple Linear ModelFitting a Multiple Linear ModelBringing Qualitative Features into the ModelModel AssessmentSummaryPart VI. ClassificationChapter 19. ClassificationExample: Wind-Damaged TreesModeling and ClassificationA Constant ModelExamining the Relationship Between Size and WindthrowModeling Proportions (and Probabilities)A Logistic ModelLog OddsUsing a Logistic CurveA Loss Function for the Logistic ModelFrom Probabilities to ClassificationThe Confusion MatrixPrecision Versus RecallSummaryChapter 20. Numerical OptimizationGradient Descent BasicsMinimizing Huber LossConvex and Differentiable Loss FunctionsVariants of Gradient DescentStochastic Gradient DescentMini-Batch Gradient DescentNewton’s MethodSummaryChapter 21. Case Study: Detecting Fake NewsQuestion and ScopeObtaining and Wrangling the DataExploring the DataExploring the PublishersExploring Publication DateExploring Words in ArticlesModelingA Single-Word ModelMultiple-Word ModelPredicting with the tf-idf TransformSummaryBibliographyData SourcesIndexAbout the AuthorsColophon

Описание

Ниже — практический обзор по теме «data».

And you want the skills required to distill a messy pile of data into actionable insights. As an aspiring data scientist, you appreciate why organizations rely on data for important decisions—whether it's for companies designing websites, cities deciding how to improve services, or scientists discovering how to stop the spread of disease. We call this the data science lifecycle: the process of collecting, wrangling, analyzing, and drawing conclusions from data.

It's aimed at those who wish to become data scientists or who already work with data scientists, and at data analysts who wish to cross the "technical/nontechnical" divide. Learning Data Science is the first book to cover foundational skills in both programming and statistics that encompass this entire lifecycle. If you have a basic knowledge of Python programming, you'll learn how to work with data using industry-standard tools like pandas.

Refine a question of interest to one that can be studied with dataPursue data collection that may involve text processing, web scraping, etc.Glean valuable insights about data through data cleaning, exploration, and visualizationLearn how to use modeling to describe the dataGeneralize findings beyond the data

Файл доступен для загрузки ниже.

data science scientists that learning wrangling exploration modeling

Частые вопросы

Можно ли скачать «Learning Data Science: Data Wrangling, Exploration, Visualization, and Modeling with Python» бесплатно?

Да, «Learning Data Science: Data Wrangling, Exploration, Visualization, and Modeling with Python» доступна для бесплатного скачивания на нашем сайте в формате PDF. Ссылка на файл находится на этой странице.

В каком формате и какого размера файл?

Книга предоставляется в формате PDF, размер файла 5,6 МБ.

Кто автор и когда вышла книга?

автор — Gonzalez Joseph , Lau Sam , Nolan Deborah (Deb), издательство O’Reilly Media, Inc., год выпуска 2023, 597 страниц.

О чём книга «Learning Data Science: Data Wrangling, Exploration, Visualization, and Modeling with Python»?

As an aspiring data scientist, you appreciate why organizations rely on data for important decisions—whether it's for companies designing websites, cities deciding how to improve services, or scientists discovering how to stop the spread of

Похожие материалы