R
R is a language and environment for statistical computing, data analysis and graphics. The competency combines reliable data manipulation with statistical modeling, package-based workflows and reproducible reporting, while keeping the interpretation of a model separate from the fact that an R function successfully fitted it.
What it is
R centers many operations on vectors, matrices, lists and data frames. Vectorized expressions, indexing and recycling rules determine how calculations apply across observations. Statistical functions often use formulas to specify responses, predictors and interactions, with fitted objects exposing estimates, diagnostics and predictions. Packages extend these capabilities, but package conventions can differ from base R. Missing values, categorical factors and date classes carry meaningful behavior. An R analysis therefore involves both programming semantics and the statistical assumptions represented by the chosen estimator.
What the work involves
A practitioner checks data classes and missingness, constructs explicit transformations and separates exploratory scripts from reusable functions. They choose a model appropriate to the observation structure, inspect diagnostics and communicate uncertainty with tables or graphics. They record package versions and make reports executable from a known dataset rather than dependent on workspace history. A useful result is a traceable analytical pipeline in which another person can reconstruct the data preparation, fitted specification and conclusions, including the limitations of the evidence.
Illustrative example
An analyst compares weekly demand across store formats. They parse dates, make store format an explicitly leveled factor and fit a model with a justified seasonal term. Residual plots and held-out weeks expose systematic errors around closures. The analyst revises the specification, reruns the script from a clean session and produces a report that links each figure to the same transformed table.
Limits and common mistakes
Recycling and implicit coercion can produce unintended calculations without an obvious failure. Factor reference levels affect coefficient interpretation, and dropping missing rows can alter the population studied. A fitted model is not evidence of valid causal inference or appropriate uncertainty. Check dimensions, contrasts, residual structure and reproducibility. R and Python overlap in applications, but familiarity with one does not establish understanding of the other's data semantics.
Prerequisites
Sources and further reading
- R manuals
Provides the official language, data manipulation and statistical environment references.
Last updated: 2026-10-10