University of Groningen, Faculty of Economics and Business

Research working group
Synthesis updated 5 October 2026

AI and the future of research at FEB

Working report for the Faculty of Economics and Business, University of Groningen, synthesis updated 5 October 2026

Summary

FEB’s research aims provide a clear starting point for its plan for AI in research. The faculty wants to sustain high-quality research and strengthen its societal impact while protecting the time available for research. [1] AI could help researchers pursue questions that were previously too demanding, examine competing explanations more thoroughly, or spend less time preparing data and other material. Whether such help improves research depends on how researchers use AI and how they assess its output.

This report examines how the use and assessment of AI can differ along three dimensions. The research cycle distinguishes activities such as developing a question, obtaining evidence, analysing results, and writing a paper. Research form distinguishes Quantitative, Qualitative, and Theory/Review work, whose methods require different checks. Researcher position, from doctoral candidate to full professor, identifies differences in learning, independent research, and supervision. These dimensions are connected, because reviewing an AI-generated analysis, for example, has a different purpose for a doctoral candidate learning the method than for a professor supervising several projects.

The most plausible gains concern work that researchers can check against records, sources or a specified procedure. Delegating the main analysis carries greater learning risks for PhD candidates, whereas preparing and validating data may benefit them directly. Researchers with relevant expertise may gain time, but those responsible for several projects can also receive more generated work to check. The combined assessment therefore differentiates tasks and responsibilities without assuming that all doctoral uses are harmful or all senior uses beneficial. These are conditional implications for FEB, not measured effects among its researchers.

FEB should therefore begin by supporting AI use in a limited set of research projects across methods and career positions. The participating teams should evaluate the quality of the completed work, the total effort including checking and supervision, and what researchers can explain independently afterwards. FEB’s working group on AI in research can then use those evaluations to recommend shared tools and support, identify where doctoral training needs a different approach, and judge whether faster work leaves time for better research.

Introduction

Researchers at FEB can increasingly ask AI systems to search papers, write analysis code, comment on a manuscript, or prepare a grant proposal. The faculty’s September 2026 strategic plan calls for a research AI plan in 2026–2027 and recognises that training and tools will cost time and money before any efficiency gains are realised. [1] The practical question is which uses of AI the faculty should support, because the same assistance may improve one part of a project but make another harder to assess.

To answer that question, the report begins with FEB’s research aims and the evidence on how research practice is changing. Its main chapters then examine the three dimensions, namely the research cycle, research form, and researcher position, and explain which work AI may change and why the consequences could differ. The final chapters bring the dimensions together in scenarios for FEB and recommendations for the working group. Throughout, the report asks how AI could improve research and where researchers need support so that AI use does not weaken the evidence for their claims, their understanding of their work, or doctoral training.

1. Research at FEB and the development of AI

Research quality and impact

The strategic plan describes what FEB means by strong research and how the faculty intends to measure it. FEB wants to remain a leading international research institution in economics and business in Europe and to obtain an excellent result in the 2027 national research evaluation. Its indicators of high-quality research are the average research output per FTE, the average number of publications per FTE in top economics and business journals, and the research evaluation itself, and funding from prestigious grants should rise each year to 20% above the 2023–2025 average by 2031. Even if staff numbers fall, the plan seeks to keep research and impact time per research-active employee broadly stable and to maintain the average quantity and quality of output per researcher. [1] These aims and indicators raise the question of how researchers will use any time that AI saves. A researcher might examine a more demanding explanation or develop a distinctive dataset, but could also be expected to produce more papers before FEB knows what it costs to check AI-assisted work.

The plan also recognises contributions whose value is not captured by publication volume. Societal impact, for example, concerns changes in business and management, public policy, and society, and preparing a competitive grant can reduce near-term publication output but support later research. [1] FEB should therefore assess an AI programme by the research it helps a team undertake and the use others can make of that research. A grant application or report that AI has polished offers little benefit if the proposed study is unconvincing or if the report’s findings do not address the intended problem.

In practice, FEB recognises publications in top journals by the standing of the journal. The plan does not define its top-journal indicator in terms of a metric, but FEB’s career and research-time rules use the Article Influence Score Percentile (AIP), which locates a journal’s Article Influence Score within the combined SSCI and SCIE journal distribution as a moving five-year average. [2] For new hires on the Research Profile, the Career Tracks document treats publications in journals with an AIP of at least 95 as indicative of high-quality work in specified assistant-to-associate promotion routes, alongside originality, rigour, grants, and other contributions. [3] FEBRI also uses (AIP/100)² as the journal-score component of the publication credits that help determine research time, so that these credits rise steeply with the standing of the journal. [2] AIP describes the journal, however, and cannot establish whether an individual paper uses a valid measure, identifies the effect it claims, or develops an original explanation. Citation performance also varies considerably between articles in the same journal. [4]

The journals themselves ask authors to make different kinds of contribution. The American Economic Review seeks work of broad interest to economists, [5] and Econometrica includes theoretical and empirical research with rigorous quantitative reasoning, [6] whereas the Journal of Economic Literature guides readers through selected bodies of economic scholarship. [7] In management, the Academy of Management Journal publishes empirical work across methods, [8] the Academy of Management Review centres on conceptual theory, [9] and the Academy of Management Annals publishes integrative reviews. [10] These differences between journals help explain why researchers cannot judge AI’s contribution to a paper from the quality of its prose or the success of its computations alone.

Research quality Possible AI assistance What researchers need to establish
Question and argument Compare questions and alternative explanations. The study adds to what readers understand.
Evidence and design Find sources, draft measures, and compare designs. The evidence and design support the intended claim.
Analysis Generate code, simulations, and additional checks. The measures, assumptions, and calculations are sound.
Interpretation and reuse Clarify explanations and document decisions. Conclusions follow from the evidence, and others can inspect the work.

Table 1. Research standards for assessing AI assistance. The middle column lists possible uses for FEB projects to evaluate. [11] [12] [13]

AI in FEB’s strategic plan

The strategic plan treats AI as a disruptive technology that affects what FEB studies as well as how its researchers work. For research, the plan expects AI to make the Digital and AI theme more prevalent and to change what is studied in disciplinary fields, and FEB aims to be at the forefront of these developments, partly by benefiting from investments in the AI Factory in Groningen. The plan also asks the research organisation to consider carefully how AI can support the execution of research, and it gives the research directors the task of developing a plan, together with internal experts and front-runners, on how FEB should use AI effectively, with a completed plan as its indicator for 2026–2027. [1]

The plan connects these ambitions to investment and responsibility. FEB intends to invest in the AI competencies of its staff together with UG, and it expects the costs of training and tools to precede any long-term efficiency gains. The plan also names responsible and ethical use, digital sovereignty, energy consumption, and the risk that AI increases existing inequalities as questions that the faculty needs to address. Its expectations of lower workload and cost savings concern mainly education and support, for example AI-assisted grading and pilots in student support and admissions. [1] For research, the more pressing question is therefore how AI can contribute to the quality and impact that the plan values.

Researchers could use AI to contribute to each of these criteria in a different way. For publications in top journals, AI may help a team test alternative explanations, run robustness checks that it could not otherwise afford, and search a literature that spans economics and several business fields, provided that the researchers check the generated code and the completeness of the sources. [14] [12] [15] For grant applications, AI can help researchers compare a draft with a funder’s call and develop alternative designs, although a proposal becomes stronger only when its design and plan improve. For societal impact, AI may help researchers translate findings for partners and analyse partner documents within the data agreement, but impact still depends on whether the findings address the partner’s problem. The chapters that follow examine where such contributions are likely and where researchers need support to realise them.

Scope across departments

AI is relevant across FEB’s seven departments, although its effects need not be equally large. [16] Publication examples show that methods vary within departments as well as between them. Marketing researchers use both synthetic data and focus groups, [17] [18] and operations researchers develop formal optimisation models and qualitative studies of logistics. [19] [20] The examples come from the publication records of current staff. Most of these papers have not been classified by method, and the examples are a small purposive sample. They support involving all departments but cannot support a departmental ranking of AI opportunities or risks. The report therefore compares research activities, forms, and positions across the faculty.

Earlier research software and current AI

Across these research activities, AI may change both the time a task requires and the questions researchers can investigate, much as earlier research software did. Statistical packages broadened access to estimation and simulation, and computers allowed economists to examine questions that would have been prohibitively expensive to calculate by hand. [21] [22] Researchers still had to understand the computation, however, and concerns about the numerical reliability of econometric software were documented in the 1990s. [23] Generative AI extends this development to language and reasoning tasks, and AI agents can search files, run code, and revise their work in response to instructions.

The growth of the literature makes such assistance attractive. In Thelwall and Sud’s historical Scopus data, the number of indexed journal articles rises from approximately 111,000 in 1950 to 2.57 million in 2020. [24] Changes in database coverage explain part of this increase, so the series cannot establish how much computing caused publication to grow. The figures nevertheless show the scale of the literature researchers now have to search and read. Assistance with finding and comparing papers could be valuable, particularly when a question spans economics and several business disciplines.

Annual Scopus-indexed journal articles and changing abstract coverage, 1900–2020

Figure 1. Annual journal-article records in the historical Scopus data. The series reflects changes in database coverage as well as publication activity. [24]

Evidence from knowledge work shows, however, that the benefit depends on the task and the user’s ability to check the result. In Dell’Acqua and colleagues’ experiment, consultants using AI worked faster and produced higher-rated work on tasks within the tested system’s capability, but were less likely to reach the correct answer on a task outside it. [25] A coding study estimated, with substantial uncertainty, a time saving in one professional engineering setting, [26] whereas another found that experienced maintainers working on their own repositories completed tasks more slowly with AI. [27] Any comparison of research productivity therefore needs to include the time researchers spend inspecting and repairing generated work.

Current AI systems can also perform demanding scientific tasks. In a specialist-refereed mathematics exercise, seven of ten previously undisclosed problems received at least one passing solution across the four tested systems, although the report on the exercise also describes unsupported steps and misleading references. [28] A system described in Nature helped write empirical scientific software for tasks whose results could be scored. [29] These results give researchers reason to reconsider which technical tasks they can delegate, but researchers still have to assess the quality of the complete research project. A model may be solved correctly, for example, even though its assumptions do not describe the problem, and a study’s table may be reproduced even though its design cannot support a causal claim.

2. Dimension 1: The research cycle

The first dimension concerns the stage of the research cycle, that is, the recurring work through which researchers develop a question into a claim that others can examine. Researchers read and frame a problem, design a way to investigate it, obtain or generate material, analyse that material, interpret the results, and communicate and revise the work. These activities are connected and frequently revisited, and the AMJ research canvas makes the relations among question, theory, design, and evidence visible without requiring every project to take the same route. [30] For a FEB researcher, the relevant issue is how assistance with one activity affects the reasoning and evidence used in the next, because a well-executed step may improve the project but may also carry forward a mistake introduced at an earlier step.

The cycle changes with the type of work. An empirical economist may return to the research design after discovering that a proposed comparison does not identify the quantity of interest, [14] and a field researcher may adjust sampling and questions as observations reveal an unexpected process. [11] A conceptual scholar, by contrast, can revise premises after testing an implication without collecting new primary observations. [31] Authors of an integrative review obtain published studies as their evidence, [32] and authors of a statistical meta-analysis also extract and model the studies’ numerical results. [33] The six stages below therefore organise a comparison of research tasks without prescribing a sequence or assuming that every project acquires new data.

Research stage Work carried through the stage Where AI assistance can be examined
Question and theory Specify the phenomenon, concepts, and intended contribution. Alternative explanations, construct distinctions, and counterexamples.
Literature Find, select, read, and relate prior work. Discovery, screening, passage retrieval, and comparison.
Design Choose the comparison, cases, model, or review procedure. Simulation, draft instruments, feasibility checks, and alternatives.
Data and source material Obtain permitted material and prepare it for analysis. Retrieval, transcription, extraction, coding, and cleaning where applicable.
Analysis and interpretation Produce results and decide what follows from them. Executable code, formal derivation, theme comparison, and error checks.
Writing and review Explain, criticise, revise, and make the work reusable. Drafts, reviewer tracking, documentation, and source checks.

Table 2. Six recurring stages used to compare research tasks. A purely conceptual project has no empirical data-acquisition step. Reviews work with prior studies, and statistical meta-analyses also extract their results.

Questions, concepts, and literature

At the beginning of a project, AI can help a researcher explore unfamiliar terms, compare the language of neighbouring fields, or identify a paper that deserves reading. Help with terminology can be valuable when an economics concept and an organisation-studies concept describe related processes without using the same vocabulary. A researcher can also ask an assistant to compare explanations of a phenomenon and then inspect the cited work. Korinek illustrates these uses for economic research, and work on source-grounded literature systems shows that retrieval can improve answers that draw on several papers. [34] [35] An exploratory conversation can be useful even if the researcher discards some of its suggestions, but a claim in a manuscript needs a stronger source trail.

Researchers may, however, take a literature to be complete merely because it reads coherently. A study of GPT-4o citation recommendations found that the model favoured highly cited work. [15] A researcher who relies on model-suggested citations may therefore miss a smaller research stream that challenges the proposed explanation. When researchers conduct an integrative review, their search and inclusion choices are part of the method. Cronin and George explain that an integrative review makes its contribution by bringing studies into a new relation, which requires knowing which studies have been considered and what they actually establish. [32] An assistant may help screen a defined collection of studies or extract information from it, but the authors need to examine disputed inclusions and the material behind a proposed synthesis. Doctoral candidates who read the studies themselves may also develop the ability to recognise a weak inference or a missing research stream.

Questions and concepts deserve particular attention, because an early error becomes harder to see when a tool completes the later steps efficiently. A tool may propose operational variables for “trust,” write the survey items, and then analyse the responses. If the items in fact measure satisfaction, the analysis cannot support the intended claim even when the code is impeccable. In quantitative research, this mismatch is a measurement problem, and in qualitative work an early error can become an interpretive problem if the tool’s summary starts to define what participants meant. In conceptual work, a term can quietly shift its meaning across paragraphs while the prose remains fluent. The research team needs to establish the construct and the evidence required to examine it before accepting a convenient implementation.

Design and access to evidence

During design, AI can generate possible experiments, prepare simulation code, suggest interview questions, or make the consequences of a statistical assumption easier to inspect. These proposals become useful when researchers can explain why one design addresses the question better than its alternatives. Empirical economists, for example, give causal identification a central place, and structural economists ask what economic assumptions permit an observed pattern to inform a counterfactual. [14] [36] In management field research, whether a focused test or an exploratory inquiry is appropriate depends on the relation between existing theory and the phenomenon. [11] An AI-generated design may be clever and still lack the comparison, access, or contextual knowledge that the intended inference requires.

The access that a design requires depends on more than the technical ability to retrieve evidence. CBS microdata require institutional and project authorisation, individual confidentiality commitments, and work inside an approved environment. [37] In organisational fieldwork, researchers’ relationships with a site and its participants influence what they can observe and ask. [38] An agent can write a query or a participant invitation, but it cannot grant rights to the records or permission to contact people. Nor does a tool’s ability to process files establish that the researcher may upload confidential data to it. At FEB, the rules governing a company collaboration, a licensed database, or a restricted administrative source thus remain part of the research design.

Once researchers have established access, AI may also change how they collect evidence. Accounting and finance researchers, for example, can obtain public SEC filings through APIs and then process them with software. [39] Chatbot interviews go further, because they alter the interview encounter itself. In a study with 74 social scientists, a chatbot asked adaptive interview questions after researcher-led recruitment and consent, but the researchers also observed missed probes, leading questions, and interruptions. [40] A web-survey experiment with 1,800 participants found that AI follow-up questions could elicit more detailed responses, although the procedure involved trade-offs in coding and participant experience. [41] Whether chatbot interviews and AI follow-up questions produce suitable research material depends on the question and the participants, so the findings need to be tested in the settings where FEB researchers would use these procedures.

Preparation, analysis, and interpretation

Preparing material can consume a large share of a project’s time, and AI can help researchers transcribe interviews, identify passages in documents, write code to merge datasets, extract a value from filings, or classify text. Each of these activities produces an output that researchers may later treat as evidence. Levich and Knust’s accounting application illustrates the reporting and language differences that arise when ownership information is extracted from corporate documents. [42] If an AI-generated label becomes a covariate or outcome, its classification errors can bias the estimate. Econometric research on model-generated measurements develops ways to validate and adjust the analyses that use them, including using human-labelled samples to estimate and correct for measurement error. [43] [44] Researchers may be able to process more records at lower cost, but they still need to establish what the generated variables measure.

Researchers can check some generated analysis code by running it and comparing its output with known results, but two reproducibility studies illustrate the distinction between execution and assessment. Brodeur and colleagues randomly assigned 288 researchers to human-only, AI-assisted, and AI-led teams for selected quantitative social-science reproduction tasks. Human-only and AI-assisted teams reproduced the assigned results at similar rates, whereas AI-led teams did much worse with the 2024 systems and workflow. Human-only teams, however, detected more major coding errors than AI-assisted teams. [12] A June 2026 preprint evaluated newer coding agents on 221 selected tasks from 54 papers, including ten tasks for which necessary material was missing. The strongest configuration that its authors tested reached 93.4% task accuracy and 78% whole-paper accuracy over three runs. [45] Because the two studies evaluate different tools, materials, and degrees of researcher involvement, their results cannot isolate the effect of improved models, but they do give FEB reason to test the precise workflow it intends to support.

When generated code runs correctly, researchers still need to interpret its output. A reproduced table may contain a coefficient that is calculated exactly as the code instructs, although the variable was built from the wrong population or the causal comparison remains confounded. For the tasks with missing material, the later replication benchmark found that a prompt nudging an agent toward the published result reduced the agent’s ability to recognise that a task was impossible. [45] Researchers need to be able to conclude that a result cannot be reproduced, even when the published table provides an apparently clear target.

Writing, review, and reuse

Researchers often discover what they can claim while explaining a result to coauthors, seminar participants, or reviewers. AI can help them clarify a passage or organise comments, as long as they preserve the reasoning that connects the evidence to the conclusion. In a short professional-writing experiment, AI assistance reduced completion time and improved assessed writing quality, although a journal article is longer and has different evidential demands. [46] In a manuscript, a polished explanation may help readers when it makes an assumption visible, but it is harmful when it masks an unresolved construct, repeats a source’s conclusion without its qualifications, or makes exploratory work look like a prediction made in advance.

AI feedback can also make an author too comfortable with a weak argument. Experiments found that biased writing suggestions shifted users’ expressed positions on societal issues, [47] and that sycophantic responses increased participants’ conviction of being right and reduced their willingness to repair interpersonal conflicts. [48] Although these experiments concern other settings, they suggest a risk when researchers ask an assistant to judge an argument they already favour. Feedback is more useful when it identifies an unsupported inference or a contradictory source that the author can examine. Authors preparing a revise-and-resubmit should keep the original reviewer comments available so that a reassuring summary does not replace a difficult criticism.

When drafting becomes easier, work can shift to the people who read the drafts. An Organization Science editorial reports a steep rise in submissions and raises concerns about the quality of AI-supported manuscripts and reviews. [49] Although the editorial’s observation cannot establish that AI caused a change in publication quality, it points to a practical concern for FEB, namely that coauthors, supervisors, and reviewers may have less time to examine each claim if papers, revisions, and grant applications arrive more rapidly. A productive AI workflow should therefore deliver more than a clean final document and make it easier for the next reader to recover how a claim was made.

At the end of a project, researchers should leave material that another researcher can inspect or reuse where rights permit. AEA data and code rules and AMJ transparency guidance already express this expectation in different research traditions. [50] [13] For AI-assisted work, relevant records may include the model and tool version, source set, instructions that affected coding or classification, validation sample, failed checks, and human decisions. Which records a project needs depends on the task and on the journal or data agreement. A record that researchers keep during the project is more reliable than one they try to reconstruct after a reviewer asks how the result was produced.

3. Dimension 2: Research form

The second dimension concerns the form of research, because the errors a researcher needs to look for in AI-assisted work depend on the methods a study uses. Quantitative, Qualitative, and Theory/Review provide practical groups for comparing the activities and methods used at FEB. The three groups are not mutually exclusive identities for departments, and they do not imply that all work within a group follows one method. Mixed studies can cross the Quantitative and Qualitative groups, [51] and theory is built and tested in empirical work as well as in papers whose principal contribution is conceptual. [52] Statistical meta-analysis belongs in Quantitative because it extracts and combines numerical results from earlier studies, and an integrative review belongs in Theory/Review because its central work is interpreting and connecting a field of scholarship. [33]

Research group Work an assistant may help with Claim that still has to be justified
Quantitative Code, extraction, statistical analysis, simulation, measurement from documents, and numerical synthesis of prior studies Whether observations and measures represent the phenomenon, the design supports inference, and results survive relevant checks
Qualitative Transcription, retrieval, organisation, structured coding, and comparison of passages Whether interpretations preserve context and can be followed from the material, including discrepant cases
Theory/Review Literature search and synthesis, counterexamples, conceptual comparisons, formal derivation, and model solving Whether premises, source selection, and reasoning support an original and useful conclusion

Table 3. The groups distinguish research activities and the claims those activities support. A single project may use more than one row, and a successful AI output in the middle column does not by itself establish the claim in the last column.

Quantitative research

In many quantitative tasks, researchers can test directly whether an assistant has carried out the work correctly. A generated script can be run on a fixed dataset and its output compared with a known table, and simulation code can be inspected against a specified model. AI may allow researchers to examine alternative estimators or run robustness checks they could not otherwise afford. The added analyses can improve the research when they address a real weakness in the inference, because a regression executed correctly on an unsuitable measure or an unidentified comparison remains a weak basis for a claim. If running many specifications becomes cheap, researchers may also find a persuasive result without being clear about which analyses were planned and which were tried after seeing the data. Researchers who report transparently and provide reusable code help readers distinguish planned tests from subsequent exploration. [14] [50] [13]

AI can also become part of the study itself, as a conversational agent or as a source of simulated responses. In three studies with 789 human participants in total, Becker and colleagues used customised conversational agents that acted as a leader, a resistant subordinate, and a teacher in a confidentiality scenario. [53] Conversational agents may let researchers create interactive treatments that a fixed vignette cannot provide. The researcher then needs to document and evaluate the agent’s behaviour, because the agent is part of what participants experience. Synthetic participants raise the different question of whether model-generated responses can stand in for observations of people. Early work showed that language models could approximate selected survey patterns, but subsequent comparisons found plausible averages alongside insufficient variation, altered regression relations, and sensitivity to prompts or timing. [54] [55] Researchers can use simulated responses to pilot a questionnaire or explore a mechanism, but treating them as observations of consumers, employees, or investors requires validation against the particular behaviour a paper wants to explain.

Statistical meta-analysis illustrates how repetitive work and scholarly judgement are combined within Quantitative research. AI may help a research team identify candidate studies, screen abstracts, extract effects, and generate analysis code. Before a pooled estimate has meaning, however, the team must decide which studies address the same question, how each effect size was defined, whether estimates are dependent, and what should be done with conflicting operationalisations. Researchers studying AI screening have used known inclusion decisions to test prompts, and a review of generative AI in evidence synthesis finds that performance varies across tasks and systems. [56] [57] The team could likewise validate screening against a checked set of decisions and then examine ambiguous exclusions and extraction disagreements. In this way, the team may finish faster, although better coverage and more careful attention to incompatible studies may be the more important benefit.

Qualitative research

Qualitative projects also include tasks that vary in how readily researchers can check an AI output. A transcript can be compared with the audio, and passages retrieved about a specified event can be checked against the source documents. When an assistant applies a well-defined codebook, researchers can evaluate its codes against human coding decisions. Xiao and colleagues, for example, tested GPT-3 on deductive coding of 668 children’s questions using an expert codebook. The model’s agreement with expert labels was higher for question complexity (κ = .61) than for syntactic structure (κ = .38), and below agreement between experts on both tasks. [58] The results suggest a possible role for AI in preliminary coding that researchers subsequently check. Interpreting an organisation’s practices, however, requires further knowledge of its members and circumstances.

Interpretation asks a different question from classification, because the meaning of an interview statement may depend on when it was made, what a participant could safely say, and what happened elsewhere in the organisation. Braun and Clarke describe thematic analysis as a recursive process in which familiarisation with the material and the researcher’s interpretive work are part of the method. [59] If a researcher first encounters the field through a model’s summary, that summary may direct the researcher’s attention away from a contradiction or an unusual participant. That risk does not rule out every qualitative use of AI, however, and methodologists disagree about whether and how AI can fit approaches such as reflexive thematic analysis. Some see delegation as incompatible with the researcher’s situated interpretation, but others outline conditions under which AI might support inquiry without taking over its central work. [60] [61] FEB should therefore base its support on each project’s stated approach and avoid promising that AI can “do qualitative coding” in general.

In mixed research, relating the quantitative and qualitative components to each other is itself a research task. Imagine an HRM study that estimates how responses to a policy differ across teams and then interviews employees to understand why. An assistant might implement the statistical model and retrieve interview passages that mention the policy. The team still has to decide whether those passages support the proposed mechanism, suggest a different explanation, or show that the survey failed to measure an important part of the employees’ experience. Methodological fit in such a study concerns whether the question, evidence, and inference are consistent across its components. [11] The assistant may help with each component without settling how the components should be combined.

Theory and review research

Theory/Review includes several forms of work that should not be collapsed into “writing.” An economist may develop a formal model, an operations researcher may formulate and solve an optimisation problem, a management scholar may develop a conceptual explanation, and a review author may reinterpret a body of published studies. All four rely on prior knowledge, assumptions, and reasoning, but each reaches a contribution by a different route. Jaakkola distinguishes designs for conceptual papers, and Cronin and George show why an integrative review requires more than summarising the articles it includes. [31] [32] Purely conceptual and formal components can proceed without new primary data collection, whereas integrative reviews assemble published studies as their evidence. If a paper also tests its argument empirically, that empirical component requires the checks appropriate to its method.

In formal work, current AI systems can do more than produce plausible prose. The mathematics results described earlier and research on language-model systems that formulate optimisation problems show that current systems can produce derivations, code, and solutions for specified problems. [28] [62] A correct solution, however, can be uninformative if its assumptions omit a constraint that defines the real problem. Makadok’s guidance for formal theory asks authors to justify the model’s assumptions, show how its results follow, and explain what they mean. [63] In economics, a model’s implications depend on its behavioural premises and the connection between those premises and data. [36] In operations, a mathematically valid scheduling solution may fail because the organisation cannot implement a necessary decision rule. Researchers can use AI to compare formulations and look for counterexamples, but they decide which assumptions are informative and what can be learned from a result.

Authors of conceptual papers face a related risk, namely that an explanation can be well formed and still weak. An assistant can assemble familiar constructs into smooth prose, name a boundary condition, and announce a contribution. Cornelissen argues, in a critique of AI-assisted theory, that such prose can easily acquire the appearance of theory without a precise mechanism or a real change in understanding. [64] Authors may then mistake a well-written explanation for a theoretical advance. For a FEB conceptual paper, the author needs to state what happens, under which conditions, why the proposed relation follows, and what the reader would understand differently from prior work, because theoretical vocabulary alone cannot make an argument original. An assistant can be useful in this work when it challenges an implication or locates a counterexample.

An integrative review has its own evidence problem, because the published studies it assembles are its evidence. Review authors decide which conversations belong together, assess differences in concepts and methods, and explain what becomes visible when the conversations are integrated. Cronin and George, as well as Elsbach and van Knippenberg, describe reviews whose new perspective remains grounded in a representative and fair treatment of the field. [32] [65] Source-grounded AI can help authors map a large corpus and compare passages, provided that their searches, exclusions, and claim-to-source links remain available. AI may also favour prominent or easily retrievable work and give the authors a falsely complete view of a field. [15] [66] Review authors need to examine what each paper contributes, where the literature conflicts, and whether the proposed synthesis follows from those contributions and conflicts. For reviews across economics and business, the benefit of faster synthesis depends on whether the authors can defend their treatment of the literature.

4. Dimension 3: Researcher position

The third dimension concerns the researcher’s position, because the same activity can serve different purposes at different points in a career. Research produces a paper, dataset, model, or explanation, and it also develops the people who will conduct later projects. For a doctoral candidate, writing and debugging code may be part of learning to identify a flawed design. When a full professor reviews a similar analysis, that review may be part of supervising a candidate and safeguarding a wider programme. Responsibilities of this kind are described in the profiles of UFO, the university job-classification system that RUG uses, [67] and FEB’s Career Tracks criteria add specific advancement conditions for the new-hire tracks. The profiles describe generic responsibilities, but actual duties, time allocation, and AI expertise vary between colleagues. [68] [3]

Position Research responsibility in the relevant UFO profile Consequence to examine when AI assists
PhD candidate Plans and conducts research under supervision, publishes, and prepares a dissertation. [69] The project must progress while the candidate develops the ability to choose, explain, and defend research decisions.
Postdoc The Onderzoeker profile covers planning, conducting, and publishing research; independence, acquisition, coordination, and guidance vary by level and appointment. [70] AI may add capacity within a fixed project or contract, but validation and an identifiable contribution still require time.
Assistant professor The Universitair docent profile combines research and teaching, with greater independence, acquisition, and guidance at UD1. [71] Assisted output interacts with the development of a research line, teaching commitments, and relevant promotion criteria.
Associate professor The Universitair hoofddocent profile includes project development, acquisition, research, coordination, and guidance. [72] AI may assist their own work and also change the checking and coordination needed across projects.
Full professor The Hoogleraar profile includes leadership of a chair’s programme, resources, research quality, staff development, and doctoral supervision. [73] An increase in assisted work can change the demands of programme leadership and mentoring.

Table 4. Generic UFO responsibilities and report-level implications. “Postdoc” is an appointment label rather than one UFO grade, and the duties actually assigned to a person may differ from the full profile. [70] [68]

Doctoral researchers

A doctoral project is both a research contribution and a period of apprenticeship. A candidate who uses AI to obtain a correct script, a literature summary, or a clean revision may advance the project. The effect on learning, however, depends on whether the candidate also comes to understand why the script is correct, which studies were excluded, and how the revised argument answers a criticism. Experiments in other learning settings have measured assisted performance and independent performance separately. In Shen and Tamkin’s randomized study, participants who learned an unfamiliar programming library with AI assistance performed worse on a later test of concepts and debugging than participants who learned without it. [74] In a school mathematics experiment, unrestricted GPT assistance improved students’ performance during practice but reduced their later unaided performance, and a more carefully designed tutor largely avoided that reduction. [75] Both experiments thus show that a good assisted product is insufficient evidence of acquired competence, although neither follows FEB doctoral candidates through a dissertation or implies that every use of AI is educationally harmful.

A three-month randomized study of patent lawyers, reported in a September 2026 NBER working paper, offers a closer professional comparison. Autor and colleagues randomly assigned 133 lawyers in eleven firms to a group with access to a custom drafting tool or to a control group. In blind ratings, the AI-assisted drafts were judged better, and the improvement was larger among junior lawyers. At an unassisted endline task completed by 91 participants, senior lawyers who had used AI edited a flawed patent more effectively than senior controls, but junior lawyers showed no average improvement, and their outcomes spread in both directions. [76] In this study, senior meant at least seven years in patent law, not an academic rank. Because the study had no unaided baseline test and lost participants before the endline, it cannot establish that a particular junior lawyer lost skill or that an experienced professor will gain it. Its results nevertheless suggest that assistance can raise the quality of the immediate output and that later independent judgement can change differently from one person to the next.

Supervisors can make the difference between assisted output and independent judgement visible in a project’s ordinary work. A candidate who uses an agent to build an estimator might explain the target quantity and diagnose an intentionally altered variable before including the result in a paper. A candidate using AI for a review might read the studies defining the review’s boundaries and explain why two apparently similar effects should not be pooled. Some practice without assistance may be necessary to develop that judgement, but other repetitive work may turn out to add little to learning. In their discussion of AI and scientific training, Messeri and Crockett explicitly leave open which routine tasks trainees need to perform themselves. [77] The division between assisted and unassisted work should therefore follow the candidate’s current knowledge and the project’s learning objective, and the supervisor can observe whether the candidate is able to defend consequential decisions after using the tool.

Postdocs and assistant professors

Postdocs and assistant professors often have to establish a recognisable contribution within a limited period, although their appointments, teaching loads, and promotion arrangements differ. This time pressure is part of the broader insecurity and progression difficulties in research careers that the OECD documents. [78] With AI, a postdoc may be able to analyse a larger document corpus or test an additional mechanism, but only if the contract leaves room to validate the measures that the AI generates. Methods research explains why this validation is substantive research work and not an administrative delay. [44] A project’s deadline and budget therefore need to allow time to validate any AI-generated measures.

Assistant professors may use AI to attempt a wider set of projects and to spend less time preparing proposals or revisions. The same assistance can, however, obscure whether a scholar has built an independent research line when many outputs look polished but the scholar did not make the central choices in them. For promotion, FEB’s new-hire Research Profile evaluates originality, significance, and rigour alongside its publication route and other criteria. The standard allocation in the Research Profile combines research and teaching, and the Education Profile has a different balance and research standard. [3] Whether faster work leaves time for a new study therefore depends on the researcher’s workload and the expectations attached to their appointment.

AI may also help an early-career scholar work across disciplinary boundaries or develop more candidate designs, and an experiment in product-innovation work gives a reason to explore that possibility. In that experiment, an individual with AI could reach quality comparable to a pair working without it, and AI helped more clearly with generating ideas than with choosing the best idea. [79] The result does not show that colleagues, mentors, or reviewers can be replaced, and it does not establish how a tenure committee would judge assisted work. A faculty programme would need to keep the scholar’s intellectual decisions visible and to make any time saved available for those decisions.

Associate and full professors

Associate and full professors still develop their own research, but their UFO profiles also place acquisition, coordination, quality, and guidance within their roles, and the full professor carries responsibility for a chair’s programme and staff development. [72] [73] AI can help a senior researcher prepare feedback, compare a grant draft with a call, or explore alternative analyses in an established programme. With domain experience, a senior researcher may find it easier to notice when an apparently plausible output conflicts with a dataset’s history, a field relationship, or a theoretical debate. The patent-law experiment shows one setting in which experienced professionals who had used AI performed better on a later unaided assessment than experienced controls. [76] The study did not show, however, that senior lawyers completed assisted tasks faster, and it does not establish a general senior advantage in academia.

As the amount of assisted work grows, senior researchers may also need to exercise more judgement. A supervisor who receives several new analysis variants between meetings may have to reconstruct which specification answers the question, whether the candidate understands the choices, and why a result was set aside. AI can help the supervisor organise that review, but the conversation in which a researcher explains a difficult decision remains a source of learning for both people. Reif and Cummings theorise that workers who turn to AI instead of colleagues for feedback may weaken the exchange through which a team learns who knows what. [80] Baygi and Huysman make a related argument about the small consultations that keep expertise circulating in organisations. [81] Both arguments suggest that FEB should examine whether AI changes how often colleagues discuss research problems and learn from one another.

Through practice, researchers also learn things that are difficult to convey in a finished paper or a list of instructions. In Collins’s study of laboratories building a new laser, usable know-how often travelled through personal contact, and previous practical experience helped researchers recognise consequential details that written descriptions had not conveyed. [82] Although the study concerns one physics technique, it suggests why a written description may be insufficient for someone who has never performed the task. At FEB, practical knowledge of this kind might involve knowing why an archival measure changed, which organisational events shaped an interview, or where a formal assumption ceases to be plausible. This knowledge could help researchers identify an AI error that would be difficult to recognise from the output alone. Experience gained before widespread AI use may provide that knowledge, but being senior does not establish that a researcher has it for the task in question.

Prior experience cannot be inferred from title alone, because a doctoral candidate may be the person most familiar with a new method or dataset and a professor may be new to coding agents. In radiology studies, years of practice and prior AI familiarity did not reliably identify who benefited from AI advice, and incorrect advice could worsen performance. [83] In a separate prediction experiment, better calibrated beliefs about one’s own ability were associated with greater gains from AI beyond baseline ability. [84] Although neither setting involves academic research, this evidence cautions against assuming that the benefit of AI varies simply with rank. For FEB, it is more useful to ask whether researchers know the specific task, are aware of the limits of that knowledge, and can obtain a second view than whether they are senior.

For a comparison of AI use at FEB, the five positions can nevertheless be grouped by their main research responsibilities. PhD candidates have an explicit training purpose, postdocs and assistant professors face distinct but often time-bound expectations for independent work, and associate and full professors have wider coordination and supervisory duties. These groups guide the comparison but leave room for differences in duties, knowledge, and pressures within each group. For any one group, the likely balance between benefit and harm depends on the task, research method, researcher’s prior knowledge, available support, and who must check the result.

5. Combined effects and possible developments at FEB

AI’s effects on research at FEB depend on how the stage of a project, its research form and the researcher’s responsibilities combine. Preparing a dataset can make a study feasible while leaving the author responsible for validating the measure. Delegating the main analysis changes more of the intellectual work and, for a PhD candidate, may also remove practice in the method the dissertation is meant to teach. The separate dimension reviews support these mechanisms. Table 5 brings them together as a proposed assessment for FEB.

Where benefits are most plausible

The clearest opportunities concern tasks whose output can be checked against sources or a specified procedure. OpenScholar demonstrates source-grounded answers to scientific questions. [35] In six clinical reviews, AI-assisted extraction followed by human checking saved a median 41 minutes per study, although one review took longer. [85] Separate extraction tests found that accuracy varied considerably between the fields extracted. [86] At FEB, such assistance could make a larger panel, a new text-based measure or a cross-country comparison feasible. Economics and business applications of text-based measurement show why its value depends on retaining the meaning of the variables and checking whether errors affect the inference. [42] [43] [87]

The balance is less clear when AI supplies the central interpretation. Code can run successfully while implementing an inappropriate comparison; coding agreement can coexist with a weak reading of an interview; and a review can sound coherent while excluding an opposing literature. Qualitative-methods and theory-development critiques explain why these failures may be difficult to see in a finished manuscript. [59] [64] Theory/Review is therefore highly exposed to AI assistance, but exposure alone does not establish a negative effect. Checked source preparation and exposition may improve the work even when the final contribution requires considerable author judgment.

Career stage adds a question that task performance alone cannot answer: is the assisted step something this researcher needs to learn to do independently? In a one-hour programming experiment, AI access lowered scores on an immediate unaided quiz. Observed differences between ways of using AI were exploratory and did not establish that a particular style caused better learning. [74] This finding supports concern about delegating a dissertation’s main analysis before the candidate can diagnose its errors. It does not imply that candidates should avoid AI throughout the cycle. A candidate can benefit from preparing and checking records, or from comparing alternative questions, while retaining practice in the consequential choices.

Senior researchers face a different constraint: assistance to one author can increase the work that others must review. The UFO profiles combine a professor’s own scholarship with supervision and responsibility for research quality. [73] A faster draft can therefore save personal writing time while adding to the flow of manuscripts and analyses that require checking across a team. The Organization Science editorial’s concern about faster production without correspondingly better research is relevant here, although its evidence does not measure supervisory workload. [49] Experience can help when it concerns the particular task; the profile title alone does not establish that expertise.

A proposed assessment across the three dimensions

Table 5 summarizes the expected balance under ordinary supervision. Opportunity means a plausible net benefit after checking and learning costs; Risk means the principal research or learning threat is likely to outweigh that gain. Mixed identifies a consequential trade-off or a broad task that includes both favourable and unfavourable uses. These are provisional judgments for specified tasks, rather than estimated effects on FEB researchers. Each task includes researcher checking and stays the same across the three cohorts.

Researcher / research form Question Literature Design Data / sources Analysis Writing / review
PhD candidates · Quantitative Mixed Mixed Mixed Opportunity Risk Mixed
PhD candidates · Qualitative Mixed Mixed Mixed Mixed Risk Mixed
PhD candidates · Theory/Review Mixed Mixed Risk Mixed Risk Mixed
Postdocs / assistant professors · Quantitative Opportunity Opportunity Mixed Opportunity Mixed Opportunity
Postdocs / assistant professors · Qualitative Opportunity Mixed Mixed Mixed Mixed Mixed
Postdocs / assistant professors · Theory/Review Opportunity Mixed Mixed Opportunity Mixed Opportunity
Associate / full professors · Quantitative Opportunity Opportunity Mixed Opportunity Mixed Mixed
Associate / full professors · Qualitative Opportunity Mixed Mixed Mixed Mixed Mixed
Associate / full professors · Theory/Review Mixed Mixed Mixed Opportunity Mixed Mixed

Table 5. Proposed directions for substantive AI assistance under ordinary supervision. Theory/Review data work concerns review corpora; pure theory has no empirical data-acquisition task in that column. The background synthesis gives the task, rationale, sources, support and sensitivity for every cell.

Some differences require explanation. Qualitative data work combines useful transcription and organisation with possible AI-assisted interviewing. In a selected panel, stated willingness was similar for an ordinary survey and an AI text interview. [88] A much smaller direct comparison found less interest in repeating an AI interview. [89] These findings leave participation and disclosure in other settings unresolved. If AI changes who participates or what people disclose, checking the transcript cannot recover the missing evidence. Literature work in Theory/Review also deserves particular scrutiny because the selected literature supplies the evidence for the contribution. An evaluation of a review tool found limited retrieval coverage in its tested reviews, while successful source-grounded question answering addresses a narrower task than establishing comprehensive coverage. [90] [35]

The proposal weighs benefits against checking and learning costs for each task, with confidence recorded separately from direction: a use can offer a plausible net benefit even though evidence about its transfer to FEB remains limited. Actual projects can also reverse a cell’s direction. A PhD candidate independently replicating a prespecified estimate faces a different learning situation from one accepting the first analysis generated by an agent.

Without a faculty AI programme

Ordinary supervision and existing university services would continue. Teams with relevant expertise and workable checking procedures could already benefit from AI. Others might complete tasks faster while accumulating uncertain measures, unexamined sources or analyses that a colleague must later reconstruct. Whether this variation currently follows departments, resources or individual skills has not been established at FEB.

The main concern is that faster production changes what people have time to learn and review. Candidates may receive answers before attempting the underlying reasoning. Postdocs and assistant professors may be expected to process more material without additional validation time. Supervisors may receive more polished drafts whose unresolved problems are harder to recognise. These are conditional developments to investigate; the research does not establish their prevalence at the faculty.

With a faculty AI support programme

Support should change the task in a way that explains the expected improvement. For extraction, this may involve a suitable human-coded validation sample and advice on measurement error. For literature work, it may involve searching for known studies and auditing exclusions. For doctoral analysis, it may involve protected practice followed by independent reconstruction of a central result. Tutoring research shows that the design of assistance can change learning outcomes, providing a reason to distinguish guided use from unrestricted answer generation. [75] [91]

The background synthesis specifies an additional intervention for each cell. Under those conditions, more tasks plausibly become beneficial, especially source preparation, checked analysis and drafting from a researcher-owned argument. The favourable directions assume that the full checking and support cost remains proportionate to the benefit. FEB would need to measure that cost in the pilots. A licence or general training session alone would not establish the proposed effects.

Several uses remain mixed even with support. An interview pilot may reveal a participation problem without resolving it, and source checking cannot by itself establish an original theoretical contribution. Candidates still need practice in the main analysis. FEB can therefore expect a support programme to improve some working arrangements while continuing to examine difficult uses individually. FEB would still need to test whether staff can perform the checks, candidates retain practice in the central method, and the resulting studies are stronger.

6. Recommendations for FEB

The working group should use the 2026–2027 planning period to evaluate a small number of existing research activities with willing teams. [1] The selected projects should together cover different research forms and career positions across the departments. Suitable starting activities include validated data preparation, source-linked literature search, reproducible analysis and manuscript revision. Include doctoral candidates in both the promising tasks and a smaller set of demanding analysis tasks so that the evaluation covers independent learning as well as output. Projects using AI-generated measures need explicit validation, and participant-facing agents need a separate assessment of recruitment, disclosure and the evidence elicited. The aim is to learn what a useful service requires before FEB commits to broad provision or to higher output expectations.

Provide support within research projects

Researchers need help while a project is under way, when a question about a measure, data agreement, or tool can still be resolved. FEB can connect existing services to those projects through a clear contact route. Methods advisers can help a team define a validation sample or an independent check, and supervisors can identify which activities candidates need to practise themselves. The Library and university AI, data, and privacy specialists can help with source access and suitable tool environments. Table 6 sets out this support as four forms of provision, which the working group could discuss with the relevant services.

Support What FEB could organise What to learn from its use
Advice A contact route connecting research methods, the Research Office, the Library, and university AI, data, and privacy expertise. Which questions teams cannot resolve themselves, and who can help.
Documentation A short research guide with economics and business examples, source-checking procedures, and records of unsuccessful uses. Which procedures colleagues can reuse and when a tool change requires another evaluation.
Technical support Help configuring suitable tools for source-based work, reproducible code, and permitted data. Whether colleagues can repeat the work and inspect its sources and decisions.
Funding and time Shared access where justified, small supported trials, and time for validation and supervision. Whether the research gains justify the full cost and which groups otherwise lack access.

Table 6. Proposed faculty support, subject to agreement on responsibilities and resources.

Whatever support FEB provides, researchers can use a tool only in ways that the conditions attached to their research material permit. Public RUG guidance and service rules, together with European research guidelines, provide relevant starting points. [92] [93] The September RUG research-guideline document is an internally circulated version whose formal adoption has not been verified. [94] An individual project may also be constrained by an ethics approval, a licensed-data agreement, or obligations to a company partner. FEB can make the route to advice clear and record answers for recurring situations, but the relevant university services resolve questions about permitted tools, storage, and data use.

Evaluate research quality, effort, and learning

In each project, the team should specify what improvement it seeks and how it will assess that improvement. For generated code, the team needs to consider whether the researcher detects its errors and can explain the model, as well as whether the code runs correctly. For AI-generated variables, a checked sample can reveal errors that change the inference. In literature work, researchers need to be able to trace claims to supporting passages and to check for omitted studies, and qualitative researchers need to explain how their interpretation follows from the material, including contradictory cases. These checks follow from the methods discussed under the three dimensions, and the team should choose them before it judges the tool’s performance.

Teams also need records of their AI-assisted work that another researcher can use, without making record-keeping a separate administrative project. A reporting checklist for the use of large language models in behavioural science recommends documenting the task, model and settings, prompts that influenced the research, data handling, human validation, and scripts. [95] Teams in the FEB pilots can adapt that checklist to the role AI plays in their study. Where a project aims to inform an organisation or policy decision, the evaluation should also cover whether the findings address the partner’s problem and can be used within the data agreement, because a polished report alone would not establish societal impact. [1]

The evaluation also needs to include the time spent on checking and supervision and to distinguish assisted performance from independent understanding. The learning and professional experiments discussed above give reason to examine these two outcomes separately. [74] [75] [76] Candidates may, for example, need to explain or repeat an essential part of the work without AI. Postdocs and assistant professors need time for validation as part of the project, and supervisors need to know whether assistance reduces their workload or mainly increases the material they must review.

By the end of the initial projects, the working group should be able to recommend which uses deserve continued support, what expertise they require, and what they cost in total. In making that recommendation, the working group should also consider who has access, because a useful procedure that depends on a private subscription or an unusually well-funded project may warrant shared provision. A procedure that saves drafting time but requires extensive additional checking may need revision. With these recommendations, FEB could invest in research benefits that the projects have demonstrated, without assuming that every task completed faster increases research capacity.

Conclusion

This report aimed to assess how AI could improve research at FEB and where researchers would need support. AI can change a substantial part of that research, but its effects depend on the stage of the project, the form of research, and the person doing the work. Quantitative researchers may gain from faster coding but still need to check measurement and identification, and qualitative researchers may save time organising material as long as they retain the close engagement needed to interpret it. Theory and review authors may explore more arguments and sources, although they may also find it harder to distinguish an original explanation from plausible text.

Researcher position adds a further consideration, because work that an experienced colleague can delegate may be precisely what a doctoral candidate needs to learn, and assistance can also increase the volume of material that supervisors must assess. In its AI plan, FEB should therefore connect practical support to research and training goals and evaluate completed work together with the effort and understanding required to produce it. The three dimensions provide a basis for deciding where AI helps the faculty’s research and where its use needs to change.

Basis of the report

This report draws on FEB documents, university and European guidance, the local academic literature collection, and checked public sources. The main review was prepared on 4 October 2026; the expanded dimension research and combined synthesis were updated on 5 October. The detailed dimension review develops the studies behind each dimension, and the proposed combined assessment records all 54 judgments and their conditions. It is a targeted synthesis, and it is neither a systematic review nor an evaluation at FEB. The May 2026 memo to the research directors informed the practical examples and support proposals. Numbered references support specific claims, and the research dossier retains source locators and verification records. The scenarios and recommendations are proposals for discussion and have not been adopted as policy.

References

References are numbered in order of first appearance. Links lead to the original publication or institutional source; local documents are linked where no public copy is registered.

↩

1. Faculty of Economics and Business, University of Groningen (2026-09). Strategic Plan Faculty of Economics and Business 2026–2031. September 2026 version; research priorities and AI plan, PDF pp. 10–11.

↩

2. FEB Research Institute, University of Groningen (2024-05-01). FEBRI criteria for output assessment. Sections “Method used”, “AIP score”, and “Corresponding research time”.

↩

3. Faculty of Economics and Business, University of Groningen (2025). FEB Career Tracks: Research Profile & Education Profile. 2025 amendment for new hires from 1 January 2026; Research Profile promotion table, PDF pp. 15–17. Criteria for other profiles and cohorts differ.

↩

4. Zhang; Rousseau; Sivertsen (2017-03-28). Science deserves to be judged by its contents not by its wrapping.

↩

5. American Economic Association (n.d.). American Economic Review Editorial Policy.

↩

6. Econometric Society; Wiley (n.d.). Econometrica Aims and Scope.

↩

7. American Economic Association (n.d.). Journal of Economic Literature Editorial Policy.

↩

8. Academy of Management (n.d.; accessed 4 October 2026). Academy of Management Journal: journal scope.

↩

9. Academy of Management (n.d.; accessed 4 October 2026). Academy of Management Review: journal scope.

↩

10. Academy of Management (n.d.; accessed 4 October 2026). Academy of Management Annals: journal scope.

↩

11. Amy C. Edmondson; Stacy E. McManus (2007-10-01). Methodological Fit in Management Field Research.

↩

12. Brodeur et al. (2026-05-28). AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science. 2024 experiment; Table 1 and limitations. The July 2026 correction concerns affiliations, not the reported results.

↩

13. Katherine A. DeCelles; Jennifer Howard-Grenville; Laszlo Tihanyi (2021). Improving the Transparency of Empirical Research Published in AMJ.

↩

14. Joshua D. Angrist; Jörn-Steffen Pischke (2010). The Credibility Revolution in Empirical Economics: How Better Research Design Is Taking the Con out of Econometrics. Research design and causal identification in empirical economics.

↩

15. Andres Algaba; Vincent Holst; Floriano Tori; Melika Mobini; Brecht Verbeken; Sylvia Wenmackers; Vincent Ginis (2025-04-03). How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?.

↩

16. Faculty of Economics and Business, University of Groningen (2025-01-06). Departments.

↩

17. Gilian R. Ponte; Jaap E. Wieringa; Tom Boot; Peter C. Verhoef (2024). Where’s Waldo? A framework for quantifying the privacy-utility trade-off in marketing applications.

↩

18. Annayah Miranda Beatrice Prosser; Lois N.M. Heung; Leda Blackwood; Saffron O’Neill; Jan Willem Bolderdijk; Tim Kurz (2024). ‘Talk amongst yourselves’: designing and evaluating a novel remotely-moderated focus group methodology for exploring group talk.

↩

19. Sezen Ece Kayacık; Beste Basciftci; Albert H. Schrotenboer; Evrim Ursavas (2025). Partially adaptive multistage stochastic programming.

↩

20. Subina Shrestha; Håvard Haarstad; Ward Rauws; Paul Buijs (2025). From experiments to organizational change: Learning from urban logistics projects in Groningen and Bergen.

↩

21. Charles G. Renfro (2004-02). Econometric Software: The First Fifty Years in Perspective.

↩

22. Roger E. Backhouse and Béatrice Cherrier (2017). ‘It’s Computers, Stupid!’ The Spread of Computers and the Changing Roles of Theoretical and Applied Economics.

↩

23. B. D. McCullough and H. D. Vinod (1999-06). The Numerical Reliability of Econometric Software.

↩

24. Mike Thelwall and Pardeep Sud (2022). Scopus 1900–2020: Growth in articles, abstracts, countries, fields, and journals. Article and supplementary annual Scopus counts; the accompanying figure is generated from the supplement.

↩

25. Fabrizio Dell’Acqua; Edward McFowland III; Ethan Mollick; Hila Lifshitz; Katherine C. Kellogg; Saran Rajendran; Lisa Krayer; François Candelon; Karim R. Lakhani (2026-03-11). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Within-capability results and the task outside the frontier; Table 7.

↩

26. Elise Paradis; Kate Grey; Quinn Madison; Daye Nam; Andrew Macvean; Vahid Meimand; Nan Zhang; Ben Ferrari-Church; Satish Chandra (2024-10-16). How much does AI impact development speed? An enterprise-based randomized controlled trial. Randomised trial among 96 engineers; adjusted estimate and uncertainty.

↩

27. Joel Becker; Nate Rush; Beth Barnes; David Rein (2025-07-10). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. Study of 16 experienced maintainers and 246 tasks; early-2025 tools.

↩

28. Mohammed Abouzaid; Nikhil Srivastava; Rachel Ward; Lauren Williams (2026-06-10). First Proof Second Batch. Second Batch report, June 2026, pp. 2–5 and 7–10; seven passing problems collectively across four systems.

↩

29. Eser Aygün et al. (2026-05-19). An AI system to help scientists write expert-level empirical software.

↩

30. Sinziana Dorobantu; Marc Gruber; Davide Ravasi; Ned Wellman (2024). The AMJ Management Research Canvas: A Tool for Conducting and Reporting Empirical Research.

↩

31. Elina Jaakkola (2020-03-09). Designing conceptual articles: four approaches. Table 1 and conceptual-design approaches; illustrative, not exhaustive. The introduction uses non-empirical in an article-form convention that includes reviews/meta-analysis.

↩

32. Matthew A. Cronin; Elizabeth George (2023-01). The Why and How of the Integrative Review.

↩

33. Hannah Snyder (2019). Literature review as a research methodology: An overview and guidelines. Sections 2.1–2.2 and Table 2; review approaches and statistical meta-analysis.

↩

34. Anton Korinek (2023-12). Generative AI for Economic Research: Use Cases and Implications for Economists.

↩

35. Akari Asai et al. (2026-02-04). Synthesizing scientific literature with retrieval-augmented language models.

↩

36. Aviv Nevo; Michael D. Whinston (2010). Taking the Dogma out of Econometrics: Structural Modeling and Credible Inference.

↩

37. Statistics Netherlands (CBS) (n.d.). Microdata: Conducting your own research. CBS access overview, steps 1–4: institutional/project approval, researcher declarations and controlled environment; not an audit of an FEB agreement.

↩

38. Juliane Riese (2018-07-24). What is ‘access’ in the context of qualitative research?. Publisher abstract checked; qualitative access as an evolving relationship. Full article unavailable.

↩

39. US Securities and Exchange Commission (2024-06-06). EDGAR Application Programming Interfaces (APIs). SEC API overview and Programmatic API Access; public filings and XBRL facts, subject to access conditions.

↩

40. Tullia Jack; Alex Cooper; Lisa Flower (2026-05-27). Automating the qualitative interview? Using Gen AI chatbots in social science research. Recruitment and application, Sample, and results on interviewer behaviour. Selected sample of 74 social scientists; no general equivalence finding.

↩

41. Soubhik Barari; Jarret Angbazo; Natalie Wang; Leah M. Christian; Elizabeth Dean; Zoe Slowinski; Brandon Sepulvado (2026-08-10). AI-Assisted Conversational Interviewing: Effects on Data Quality and Respondent Experience. Survey Research Methods 20(2), pp. 161–180; web-survey experiment with 1,800 participants, not field ethnography.

↩

42. Levich; Knust (2025-12). Discriminative meets generative: Automated information retrieval from unstructured corporate documents via (large) language models.

↩

43. Christensen; Hansen (2026-05). Performing Valid Inference with AI/ML-Generated Covariates: A Guide for Empirical Practice. AI-generated covariates and inference, pp. 92–97.

↩

44. Jens Ludwig; Sendhil Mullainathan; Ashesh Rambachan (2026). Large Language Models: An Applied Econometric Framework.

↩

45. Meysam Alizadeh; Mohsen Mosleh; Fabrizio Gilardi; Atoosa Kasirzadeh; Joshua A. Tucker (2026-06-09). AI Coding Agents Can Reproduce Social Science Findings. June 2026 preprint, v1; §§2.1, 2.4, 5.1–5.3 and Table 1. Task accuracy includes missing-material tasks; selected executable results were manually reproduced beforehand.

↩

46. Shakked Noy; Whitney Zhang (2023-07-14). Experimental evidence on the productivity effects of generative artificial intelligence.

↩

47. Sterling Williams-Ceci; Maurice Jakesch; Advait Bhat; Kowe Kadoma; Lior Zalmanson; Mor Naaman (2026-03-11). Biased AI writing assistants shift users’ attitudes on societal issues. Science Advances 12(11), eadw5578. Original PDF pp. 4, 7–10: experimental results, warning conditions and limits. Societal-attitude writing tasks, not academic papers.

↩

48. Myra Cheng; Cinoo Lee; Pranav Khadpe; Sunny Yu; Dyllan Han; Dan Jurafsky (2026-03-26). Sycophantic AI decreases prosocial intentions and promotes dependence. Science 391(6792), eaec8352. Original PDF pp. 2, 5, 7–9: interpersonal conflict experiments and limitations.

↩

49. Claudine Gartenberg; Sharique Hasan; Alex Murray; Lamar Pierce (2026-04-27). More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review. Editorial analysis of submissions and review; journal-specific evidence.

↩

50. American Economic Association (2026-10-04). AEA Data and Code Policies and Guidance.

↩

51. John W. Creswell; J. David Creswell (2017-12). Research Design: Qualitative, Quantitative, and Mixed Methods Approaches. Official publisher companion overview for the fifth edition; quantitative, qualitative and mixed approaches, and experimental/survey designs. Full book not inspected.

↩

52. Jason A. Colquitt; Cindy P. Zapata-Phelan (2007). Trends in Theory Building and Theory Testing: A Five-Decade Study of the Academy of Management Journal. Figure 1 and pp. 1281–1285. Theory-building and theory-testing taxonomy for empirical AMJ articles, not all economics/business research.

↩

53. Becker et al. (2026-03-26). Using customized conversational AI agents in leadership and management research: Benefits practical illustrations and best practices. Customised conversational agents; three studies with 789 human participants.

↩

54. Argyle et al. (2023-02-21). Out of One Many: Using Language Models to Simulate Human Samples.

↩

55. Bisbee et al. (2024-05-17). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Results on response distributions, regression estimates, and sensitivity to prompt and timing.

↩

56. Tran et al. (2024-05-21). Sensitivity and Specificity of Using GPT-3.5 Turbo Models for Title and Abstract Screening in Systematic Reviews and Meta-analyses.

↩

57. Clark et al. (2025-07). Generative artificial intelligence use in evidence synthesis: A systematic review.

↩

58. Xiao; Yuan; Liao; Abdelghani; Oudeyer (2023-03). Supporting Qualitative Analysis with Large Language Models: Combining Codebook with GPT-3 for Deductive Coding. Deductive coding with an expert-specified codebook; results are task-specific.

↩

59. Virginia Braun; Victoria Clarke (2006). Using thematic analysis in psychology. Thematic analysis and recursive analytic phases, pp. 78–91.

↩

60. Jowsey; Braun; Clarke; Lupton; Fine (2025-12-17). We Reject the Use of Generative Artificial Intelligence for Reflexive Qualitative Research. Methodological position article; online publication December 2025.

↩

61. Wise; Gresalfi; Spencer-Smith (2026-03-16). Why AI is Not the Enemy: Opportunities to Strengthen Core Commitments of Qualitative Inquiry Through Trustworthy AI-in-the-Loop Analysis. Methodological position article on purposeful AI use in qualitative inquiry.

↩

62. Huang et al. (2025-05-08). ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling.

↩

63. Richard Makadok (2022-02-16). Guidance for AMR Authors about Making Formal Theory Accessible. Academy of Management Review 47(2), 193–205. Editorial guidance on formal theory.

↩

64. Joep P. Cornelissen (2026-05-01). The Artificial Intelligence Con: The commodification of theory and the new and improved credibility crisis. Organization Studies Agora argument; local accepted manuscript, original PDF pp. 3–8. Conceptual critique, not an estimated AI effect.

↩

65. Kimberly D. Elsbach; Daan van Knippenberg (2020-05-11). Creating High-Impact Literature Reviews: An Argument for 'Integrative Reviews'. Journal of Management Studies 57(6), 1277–1289. Integrative-review methodology and argument.

↩

66. Lisa Messeri; M. J. Crockett (2024-03-06). Artificial intelligence and illusions of understanding in scientific research.

↩

67. University of Groningen (n.d.). Legal protection, legislation and regulations. Section Job ranking and description; explicit use of UFO at RUG.

↩

68. Universiteiten van Nederland (n.d.). University Job Classification (UFO) manual. Manual pp. 5–6: generic profile, applicable result areas and job levels; not every activity applies to every appointment.

↩

69. Universiteiten van Nederland / UFO (2023-03-01). Promovendus. Promovendus, v11.0, 1 March 2023. Result areas 1–7, PDF pp. 2–3; level p. 5. Employee profile; doctoral statuses vary.

↩

70. Universiteiten van Nederland / UFO (2023-03-01). Onderzoeker. Onderzoeker, v11.0, 1 March 2023. Result areas pp. 2–4; level criteria pp. 7–8. Broader than postdoc; some duties are variants.

↩

71. Universiteiten van Nederland / UFO (2023-03-01). Universitair docent. Universitair docent, v11.0, 1 March 2023. Research result areas pp. 3–5; UD1/UD2 criteria p. 8.

↩

72. Universiteiten van Nederland / UFO (2021-08-01). Universitair hoofddocent. Universitair hoofddocent: directly served PDF identifies v10.0, 1 August 2021; research result areas pp. 4–7 and levels p. 9. The same URL has a different header in the search cache; verify HR version before individual classification.

↩

73. Universiteiten van Nederland / UFO (2023-03-01). Hoogleraar. Hoogleraar, v11.0, 1 March 2023. Result areas 1–3, 5, 8–10, PDF pp. 2–6; levels p. 8.

↩

74. Judy Hanwen Shen; Alex Tamkin (2026-01-28). How AI Impacts Skill Formation. January 2026 preprint; learning an unfamiliar Python library and subsequent knowledge assessment.

↩

75. Bastani et al. (2025-06-25). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Practice performance and later unaided assessment; comparison of the base and safeguarded tutor conditions.

↩

76. David Autor; Tanya Rodchenko; Josh Martin; Zanna Iscenko; Scott Strand; David Pearl; Melissa Ferere (2026-09). Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting. NBER Working Paper 35720, not peer reviewed. Original PDF pp. 2–5, 17–19, 23–24 and Tables 2 and 6; endline sample 91 of 133 lawyers.

↩

77. Lisa Messeri; M. J. Crockett (2026-05-19). The uncritical adoption of AI in science is alarming - we urgently need guard rails. Nature 653, 675–676. Comment, not a direct test of doctoral learning.

↩

78. OECD (2021-05-20). Reducing the precarity of academic research careers.

↩

79. Fabrizio Dell’Acqua; Charles Ayoubi; Hila Lifshitz; Raffaella Sadun; Ethan Mollick; Lilach Mollick; Yi Han; and colleagues (2026-06-12). The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork. Organization Science, online 12 June 2026. One-day product-innovation experiment; original PDF pp. 2–4, 12–18 and 20–22.

↩

80. Reif; Cummings (2026-07-17). The Generative AI Dilemma in Knowledge-Intensive Teams.

↩

81. Reza M. Baygi; Marleen Huysman (2025-09-24). Generative AI and the social fabric of organizations. Strategic Organization 24(2), 374–387 (2026); first online 2025. Conceptual essay; original pp. 378–382.

↩

82. H. M. Collins (1974-04). The TEA Set: Tacit Knowledge and Scientific Networks. Original journal article, pp. 177–178 and 182–184, laboratory know-how and scientific contact; one physics case, not a test of AI or academic rank.

↩

83. Feiyang Yu; Alex Moehring; Oishi Banerjee; Tobias Salz; Nikhil Agarwal; Pranav Rajpurkar (2024-03-19). Heterogeneity and predictors of the effects of AI assistance on radiologists. Results on experience-based characteristics and AI error; shares the underlying case/radiologist collection with Agarwal et al., not an independent replication.

↩

84. Andrew Caplin; David Deming; Shangwen Li; Daniel Martin; Philip Marx; Ben Weidmann; Kadachi Jiada Ye (2025-10-24). The ABCs of Who Benefits from Working with AI: Ability, Beliefs, and Calibration. Published article PDF, abstract and pp. 1–2, 5–6; controlled age-classification task, calibration conditional on baseline ability.

↩

85. Gerald Gartlehner et al. (2025-11-04). Artificial Intelligence-Assisted Data Extraction With a Large Language Model: A Study Within Reviews.

↩

86. Zuhaer Yisha; Peng Zou; Sheng Li; et al. (2026-01-14). Assessing data extraction in randomized clinical trials with large language models.

↩

87. Junting Duan; Markus Pelger (2026-07). Inference with AI-Generated Covariates.

↩

88. Felix Chopra; Ingar Haaland (2026-06). Conducting Qualitative Interviews with AI.

↩

89. Alexander Wuttke; Matthias Aßenmacher; Christopher Klamm; Max M. Lang; Quirin Würschinger; Frauke Kreuter (2025-05). AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers.

↩

90. Oscar Lau; Su Golder (2025-09-27). Comparison of Elicit AI and Traditional Literature Searching in Evidence Syntheses Using Four Case Studies.

↩

91. Greg Kestin; Kelly Miller; Anna Klales; Timothy Milbourne; Gregorio Ponti (2025-05-08). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. University physics experiment; immediate learning outcomes from a purpose-designed tutor.

↩

92. RUG AI Office (2026-09-09). FAQ - The Do's and Don'ts in Mistral-Le Chat (Vibe).

↩

93. European Commission DG Research and Innovation / ERA Forum (2026-05). Living guidelines on the responsible use of generative AI in research. Third edition, May 2026; guidance for researchers and research organisations.

↩

94. Rijksuniversiteit Groningen (2026-09-28). Richtlijn voor verantwoord gebruik van generatieve AI in de onderzoekspraktijk aan de Rijksuniversiteit Groningen. Internally shared version of 28 September 2026; formal adoption not verified.

↩

95. Stefan Feuerriegel et al. (2026-04-06). A reporting checklist for large language models in behavioural science. Nature Human Behaviour Comment; GUIDE-LLM table and qualifications, original PDF pp. 1–3.