Skip to main content
뒤로

Exam 1 Study Guide: Data Collection and Data Organization in Introductory Statistics

스터디 가이드 - 스마트 노트

자료에 맞춘 맞춤형 노트, 핵심 정의, 예시, 맥락을 확장해 제공합니다.

Chapter 1: Data Collection

Section 1.1: Introduction to the Practice of Statistics

Statistics is the science of collecting, organizing, analyzing, and interpreting data to make decisions. Understanding the foundational concepts is essential for effective statistical thinking and application.

  • Statistics: The study of methods for collecting, analyzing, interpreting, and presenting data.

  • Population: The entire group of individuals or items of interest in a study.

  • Sample: A subset of the population selected for analysis.

  • Individuals: The entities (people, objects, events) being studied.

  • Statistical Thinking: Involves understanding variability, data collection methods, and drawing conclusions based on evidence.

  • Qualitative Variable: A variable that describes qualities or categories (e.g., color, gender).

  • Quantitative Variable: A variable that represents numerical values (e.g., height, age).

  • Discrete Variable: Quantitative variable with countable values (e.g., number of students).

  • Continuous Variable: Quantitative variable with infinite possible values within a range (e.g., weight).

  • Levels of Measurement: Four types: nominal, ordinal, interval, ratio.

Example: Gender is a qualitative variable (nominal level); height is a quantitative, continuous variable (ratio level).

Section 1.2: Observational Studies Versus Designed Experiments

Understanding the distinction between observational studies and designed experiments is crucial for interpreting results and drawing valid conclusions.

  • Observational Study: Researchers observe subjects without intervention.

  • Designed Experiment: Researchers apply treatments and observe effects.

  • Types of Observational Studies:

    • Cross-sectional: Data collected at one point in time.

    • Case-control: Subjects with a characteristic are compared to those without.

    • Cohort: Subjects are followed over time.

Example: Studying the effect of a drug by assigning it to some patients (experiment) vs. observing patients who already take the drug (observational study).

Section 1.3: Simple Random Sampling

Simple random sampling ensures every member of the population has an equal chance of being selected, minimizing bias.

  • Simple Random Sample: A sample chosen so that every possible sample of the same size has an equal chance of being selected.

  • Determining a Simple Random Sample: Use random number generators or lottery methods.

Example: Drawing names from a hat to select participants.

Section 1.4: Other Effective Sampling Methods

Alternative sampling methods can improve efficiency or address specific research needs.

  • Stratified Sample: Population divided into subgroups (strata), then random samples taken from each.

  • Systematic Sample: Select every kth individual from a list after a random start.

  • Cluster Sample: Population divided into clusters, some clusters are randomly selected, and all individuals in chosen clusters are sampled.

Example: Stratified sampling by age group; systematic sampling by selecting every 10th person; cluster sampling by randomly choosing classrooms and surveying all students in those rooms.

Section 1.5: Bias in Sampling

Bias can distort results and lead to incorrect conclusions. Identifying and addressing bias is essential for valid statistical inference.

  • Types of Bias:

    • Sampling Bias: Sample is not representative of the population.

    • Nonresponse Bias: Selected individuals do not respond.

    • Response Bias: Responses are inaccurate or misleading.

  • Addressing Bias: Use random sampling, increase response rates, design clear questions.

Example: Surveying only morning students introduces sampling bias if afternoon students differ.

Section 1.6: The Design of Experiments

Careful experiment design ensures reliable and valid results. Different designs address various research needs and control for confounding variables.

  • Characteristics of an Experiment: Control, randomization, replication.

  • Steps in Designing an Experiment:

    1. Identify the problem.

    2. Determine factors and levels.

    3. Randomly assign subjects.

    4. Apply treatments.

    5. Collect and analyze data.

  • Completely Randomized Design: Subjects assigned to treatments purely by chance.

  • Randomized Block Design: Subjects grouped by a characteristic, then randomly assigned within blocks.

  • Matched-Pairs Design: Subjects paired based on similarity, one receives treatment, the other control.

Example: Testing a new fertilizer by randomly assigning plots (completely randomized), grouping plots by soil type (randomized block), or pairing plots by size (matched-pairs).

Chapter 2: Organizing and Summarizing Data

Section 2.1: Organizing Qualitative Data

Qualitative data is organized to reveal patterns and frequencies. Visual displays help communicate findings effectively.

  • Frequency Distribution: Table showing counts for each category.

  • Relative Frequency Distribution: Table showing proportions for each category.

  • Bar Graph: Uses bars to represent frequencies of categories.

  • Pie Chart: Shows proportions as slices of a circle.

Example: Survey results on favorite colors displayed in a bar graph.

Section 2.2: Organizing Quantitative Data: The Popular Displays

Quantitative data can be organized in tables and visualized using various graphs to understand distribution and central tendency.

  • Histogram: Graph for quantitative data; bars represent frequency for intervals.

  • Discrete Data: Data with distinct, separate values.

  • Continuous Data: Data with values in a range.

  • Dot Plot: Dots represent individual data points.

  • Shapes of Distributions: Symmetric, skewed left, skewed right, uniform, bimodal.

Example: Heights of students shown in a histogram; number of pets per student in a dot plot.

Section 2.3: Additional Displays of Quantitative Data

Advanced displays provide deeper insight into data distribution and trends over time.

  • Stem-and-Leaf Plot: Shows data distribution while retaining original values.

  • Frequency Polygon: Line graph connecting frequencies at interval midpoints.

  • Cumulative Frequency Table: Shows running total of frequencies.

  • Ogive: Graph of cumulative frequencies.

  • Time-Series Graph: Plots data points over time to show trends.

Example: Monthly sales data displayed in a time-series graph.

Section 2.4: Graphical Misrepresentations of Data

Misleading graphs can distort interpretation. Recognizing and correcting these issues is vital for accurate communication.

  • Misleading Graphs: Can result from inappropriate scales, omitted baselines, or distorted proportions.

  • Correcting Misleading Graphs: Use consistent scales, include baselines, avoid exaggeration.

Example: A bar graph with a truncated y-axis exaggerates differences between groups.

Table: Levels of Measurement

Level

Description

Example

Nominal

Categories without order

Gender, color

Ordinal

Categories with order

Class rank, satisfaction rating

Interval

Ordered, equal intervals, no true zero

Temperature (°C)

Ratio

Ordered, equal intervals, true zero

Height, weight

Table: Types of Sampling Methods

Method

Description

Example

Simple Random

Every individual has equal chance

Randomly selecting students

Stratified

Divide into strata, sample from each

Sampling by age group

Systematic

Select every kth individual

Every 10th person

Cluster

Divide into clusters, sample clusters

Randomly select classrooms

Table: Types of Bias

Type

Description

Example

Sampling Bias

Sample not representative

Surveying only morning students

Nonresponse Bias

Selected individuals do not respond

Mail survey with low response

Response Bias

Inaccurate or misleading answers

Leading questions in survey

Table: Shapes of Distributions

Shape

Description

Symmetric

Both sides mirror each other

Skewed Left

Tail extends to the left

Skewed Right

Tail extends to the right

Uniform

All values equally likely

Bimodal

Two peaks

Key Formulas

  • Relative Frequency:

Additional info: Academic context and examples were added to clarify definitions and applications for each topic.

Pearson Logo

스터디 프렙