AI-lab economic research needs a data-access constitution
Labs are hiring economists and funding outside researchers because they hold unusually rich usage data. Independence depends less on job titles than on who can ask questions, inspect transformations, publish null results and challenge the dataset.

What happened
The Financial Times documented the growth of economics teams and external research programmes at major AI labs, alongside concerns about proprietary data access, agenda-setting and conflicts of interest.
Why it matters
Evidence about AI and work can become structurally dependent on the firms whose products are being evaluated unless access, publication and replication rights are defined in advance.
The Financial Times reported on October 2 that AI labs are expanding economics teams and collaborations with outside researchers. The attraction is obvious: labs can observe product use at a scale and granularity unavailable to most public researchers. The concern is equally obvious: the data owner can shape the feasible questions, variables, sample and release process.
OpenAI describes its Economic Research Exchange as support for original, privacy-preserving independent projects. Anthropic’s Economic Futures programme funds external work and publishes aggregated Economic Index research using a privacy-preserving analysis system. Those are meaningful openings. They do not, by themselves, resolve selection, publication or replication risk.
Independence is an operating design
A credible collaboration should begin with a public data-access constitution. It should state who selects researchers, which questions are in scope, what transformations the lab performs before access, which variables are unavailable, how privacy protection changes inference, and whether the researcher can publish unfavourable or null findings without sponsor approval.
Pre-register the analysis where feasible. Preserve a versioned data dictionary and transformation log. Give an independent methods reviewer enough information to assess exclusions, missingness, classification error and model-generated labels. If raw data cannot leave the lab, provide a secure route for approved robustness checks and disclose what cannot be replicated.
Distinguish three products: internal descriptive analysis, sponsored external research and genuinely independent replication. All can be useful, but the label should tell a reader which party controlled the question, data construction, analysis and publication.
Selection deserves its own table. Report how many researchers applied, the criteria used, disciplinary and institutional mix, conflicts declared, projects rejected after data review and studies that stopped before publication. Without that denominator, a visible cohort cannot show whether the programme admits questions that challenge the sponsor’s commercial narrative.
Data construction should be challengeable too. Usage classifications may rely on model-generated labels, occupation mappings, language filters or exclusions of short and sensitive conversations. Publish validation samples and error bounds for the variables that support headline claims. If privacy rules prevent review of a subgroup, say that the subgroup is unmeasured rather than silently absorbing it into an aggregate.
Publication rights need a clock. Define the sponsor’s security and privacy review window, permitted redactions and an escalation route for disputes. The sponsor can protect users and systems without acquiring an open-ended veto over interpretation. A public register should show completed, withdrawn and delayed projects, with researcher-authored reasons where disclosure is safe.
Replication can be tiered. A public synthetic dataset can test code paths; a secure enclave can support approved checks against real aggregates; an independent auditor can verify the largest claims. None is equivalent to open raw data, so the governance appendix should state which layer was actually used.
Usage data also has a boundary problem. Product interactions show what customers did within one service, not what non-users did, how work changed outside the tool, or whether reported time savings improved productivity, job quality or distributional outcomes. Linking to surveys, administrative data or field studies may improve coverage, but each introduces its own selection and governance constraints.
The counterargument is that strict access rules will slow research and increase privacy risk. A constitution does not require open raw transcripts. It requires the lab to state the trade-offs, separate privacy review from result approval and make the largest reproducibility gap visible.
Decision-makers should therefore treat lab research as one evidence layer. Compare it with public statistics, independent surveys and studies whose data generation is not controlled by the vendor. When findings diverge, inspect populations, task definitions and exposure windows before choosing a headline.
The immediate decision is for every lab-funded economic study to publish a one-page governance appendix covering question rights, data construction, researcher access, privacy transformations, publication rights, robustness access and unresolved replication limits.