Research Article · Journal of Technology Management & Innovation

Multi-Agent Systems with Generative AI In Public Innovation Funding Programs: Architecture and Evaluation

Luiz Gustavo Ferreira da Silva Trufilho1*iD, Fábio Luís Falchi de Magalhães1iD, João de Paula Ribeiro Neto2iD

1 Postgraduate Program in Technological Innovation, Universidade Federal de São Paulo (UNIFESP), São José dos Campos, Brazil.

2 Postgraduate Program in Administration, Universidade Municipal de São Caetano do Sul (USCS), São Caetano do Sul, Brazil.

* Corresponding author: [email protected]

Vol. 21, No. 2, pp. 84–98 (2026)
License This journal and its contents are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0)
Received 14 May 2026 · Accepted 24 Jul 2026 · Published 7 Aug 2026

Abstract

Public innovation funding programs are central instruments of innovation policy across economies, with initiatives such as the Small Business Innovation Research (SBIR) in the United States, the European Innovation Council under Horizon Europe and the Centelha Program in Brazil. Despite robust instrumental designs, proposal preparation for these programs constitutes a systemic barrier for small and medium-sized enterprises. This paper describes the development and exploratory empirical evaluation of a technical-technological product in the Software category: a multi-agent system based on generative artificial intelligence, organized in five functional layers and fifteen hierarchical agents. The interface runs on WhatsApp and orchestration is handled by the low-code platform n8n. The research adopted Design Science Research and the artifact was evaluated, in an exploratory qualitative study of the Centelha Program case, through semi-structured interviews with ten specialists, analyzed using the Gioia, Corley and Hamilton method. Four aggregate dimensions and four research propositions emerged from this analysis. Results indicate adherence to the program’s official criteria and high adoption intent. The articulation between multi-agent systems, large language models and ubiquitous communication channels offers a concrete path to democratize access to public innovation funding, with potential transferability to other programs and human review safeguards to mitigate hallucinations and ethical risks identified.

Keywords: multi-agent systemsgenerative artificial intelligencelarge language modelspublic innovation funding policiesCentelha Program

Introduction

Public funding for business innovation occupies a central position in contemporary innovation policy. Edler and Fagerberg (2017) classify the instruments available to governments into direct funding of research and development, tax incentives and demand-side measures, and document their diffusion across advanced and emerging economies. Consolidated initiatives include the Small Business Innovation Research (SBIR) and Small Business Technology Transfer (STTR) programs in the United States, the European Innovation Council Accelerator and the Eurostars instrument under Horizon Europe, and the Smart Grants operated by Innovate UK (Lerner, 2010). In Brazil, direct instruments such as economic grants, subsidized financing and research scholarships coexist with tax incentives and are operated by agencies including the Funding Authority for Studies and Projects (Finep), the National Council for Scientific and Technological Development (CNPq) and state research foundations (Mazzucato, 2018; Rapini & Rocha, 2014).

The economic rationale for public intervention originates in the analysis of market failure in knowledge production. Arrow (1962) demonstrated that returns to invention are only partially appropriable by the investing firm, which leads private investment in research to fall below the social optimum. Subsequent evidence shows that this shortage is most severe at the earliest stages of technology-based ventures, in the range often described as the valley of death, and that well-designed early-stage grants raise the survival, patenting and subsequent private financing of supported firms (Howell, 2017; Lerner, 2010). Programs such as SBIR Phase I, the EIC Accelerator and the Centelha Program, created by the Brazilian Ministry of Science, Technology and Innovation (MCTI) in partnership with Finep, operate precisely in this instrumental niche (Castro et al., 2022; Mazzucato, 2018).

Effective access to these instruments is nonetheless mediated by the technical complexity of proposals. Calls require applicants to articulate innovative merit, technical feasibility, a business model and a financial plan under strict formal rules, which turns submission into a selective process with high cognitive and operational cost. Small and medium-sized enterprises, which rarely maintain dedicated fundraising teams, are disproportionately affected, and the resulting asymmetry of technical and argumentative capital restricts the diversity of applicants and territories reached by public funding systems (Edler & Fagerberg, 2017; Gama & Magistretti, 2025). Evidence from Brazil reinforces this diagnosis: even for a consolidated instrument such as the R&D tax incentives of the Lei do Bem, effective uptake by firms has remained below governmental expectations, with implementation barriers and the complexity of access among the identified causes (Fabiani & Sbragia, 2014).

In parallel, two research strands in computer science have matured. Large language models (LLMs) built on the Transformer architecture generate coherent text, analyze documents and perform reasoning tasks from a small number of examples (Brown et al., 2020; Vaswani et al., 2017). Multi-agent systems (MAS), investigated since the consolidation of distributed artificial intelligence, provide the organizational abstractions required to coordinate autonomous computational entities around shared goals (Dorri et al., 2018; Wooldridge, 2009). Recent studies document the convergence of the two strands and the application of LLM-based agents to knowledge-intensive domains such as healthcare, software engineering, recommender systems and supply chains (He et al., 2025; Ooi et al., 2025; Wang et al., 2024; Wu et al., 2024).

The literature analysis conducted in this research identified, however, an evident gap at this intersection: no study was found that applies MAS with LLMs to the complete cycle of proposal preparation for public innovation funding programs, a cycle that covers call analysis, criteria extraction, staged textual structuring and scoring simulation. Outside the academic literature, a commercial ecosystem of AI-assisted grant-writing tools has emerged, with services such as Granted AI, Grantable and Grant Assistant oriented mainly to philanthropic and nonprofit grant applications. The existence of this ecosystem corroborates the practical relevance of the problem, but these tools are typically closed, single-agent writing assistants, documented without peer-reviewed evaluation and not aligned with the phase structure and official scoring criteria of a specific public program. The contribution of this study differs in kind: it specifies and evaluates empirically, under a design science protocol, a hierarchical multi-agent architecture aligned with the official criteria of a national funding program and delivered through a ubiquitous conversational channel.

The research question is therefore formulated as follows: how can a multi-agent architecture orchestrated by generative AI automate and qualify the structuring of proposals for public innovation funding programs? The general objective of this work is to develop and evaluate empirically, in an exploratory study, a technical-technological product in the Software category, materialized as a multi-agent system based on generative AI supporting proposal preparation, taking the Centelha Program (Brazil) as the application case for its representativeness among early-stage subvention calls. The specific objectives are: (i) to characterize the functional and architectural requirements of the artifact from the official phases and criteria of the program; (ii) to design and implement the architecture in functional layers and a hierarchy of specialized agents, integrated into a ubiquitous conversational channel; (iii) to empirically evaluate the artifact through semi-structured interviews with specialists, analyzed using the Gioia, Corley and Hamilton method; and (iv) to formalize theoretical propositions derived from the inductive analysis, expanding the knowledge base on AI applied to innovation management. The study is justified on three grounds: theoretically, by addressing a specific gap at the intersection of MAS, LLMs and public innovation funding; practically, by offering a functional artifact with marginal operational cost; and socially, by reducing asymmetries of technical-argumentative capital among applicants in international and Brazilian programs. The methodological path followed Design Science Research and culminated in an exploratory evaluation by ten specialists (Dresch et al., 2015b; Gioia et al., 2013; Hevner et al., 2004).

The paper is organized in five additional parts. Section 2 synthesizes the theoretical framework. Section 3 details the methodology. Section 4 presents the architecture of the artifact and the results of the empirical evaluation. Section 5 discusses contributions, limitations and implications for international and Brazilian funding programs. Section 6 presents the final remarks.

Literature Review

Technological innovation and public funding

Innovation is the engine of capitalist development, described as creative destruction, a process that introduces new combinations capable of altering market structures. Technological innovation corresponds to the introduction of a new or significantly improved product, good, service or process and is a central condition of contemporary competitiveness (Schumpeter, 1934).

Public funding for innovation is justified by market failure theory and materializes in national innovation systems. The international literature on innovation policy identifies three families of instruments: direct funding of research and development, tax incentives and demand-side instruments, all articulated to correct systemic failures and direct efforts towards public missions (Edler & Fagerberg, 2017; Mazzucato, 2018). Government programs supporting early-stage innovative firms, such as the United States SBIR/STTR, the European Innovation Council Accelerator under Horizon Europe and the Innovate UK Smart Grants, are consolidated cases of this international arrangement (Lerner, 2010). In the Brazilian arrangement, the Triple Helix model operates, in which government, science and technology institutions and firms act in a coordinated manner, complemented by development banks and mechanisms such as the Lei do Bem. The main instruments include economic grants (Finep, state foundations), subsidized financing (Finep, BNDES), research scholarships (CNPq, Capes), tax incentives (MCTI) and the Centelha Program itself (MCTI/Finep), in an arrangement that combines direct and indirect resources (Gama & Magistretti, 2025; Rapini & Rocha, 2014). Analyses of the successive generations of Brazilian federal programs aimed at technology-based companies — Startup Brazil, Finep Startup and Conecta Startup Brazil, among others — show an increasing reliance on ecosystem collaboration and on formal partnerships between firms and research institutions (Bezerra Borges et al., 2021), a policy trajectory in which the Centelha Program occupies the earliest, idea-stage niche.

The Centelha Program

Taken as the central case of this study, the Centelha Program exemplifies, in the Brazilian context, the international class of early-stage subvention programs. Its structure is comparable, in purpose and evaluation format, to the United States SBIR Phase I and the Innovate UK Smart Grants, which makes the findings of this study potentially transferable to other jurisdictions.

The Centelha Program was established by Ordinance 4.082/2018 of the then Ministry of Science, Technology, Innovations and Communications and operated in partnership with Finep, CNPq, the National Council of State Research Foundations (Confap) and the CERTI Foundation. The program acts at the early stages of technology-based firms, offering economic grants, training and ecosystem connection. In its two previous editions, it was present in more than 1,500 municipalities, with over 26,500 submitted ideas (Castro et al., 2022; Ministry of Science, Technology, Innovations and Communications [MCTIC], 2018).

The selection structure is organized as a three-phase funnel of progressive rigor, illustrated in Figure 1. Phase 1 covers the Innovative Idea, Phase 2 corresponds to the Business Project and Phase 3 to the Funding Project. The scoring formulas are summarized in Table 1. The non-linear nature of the formulas, multiplicative in Phases 1 and 2 and with a risk reduction factor in Phase 2, amplifies differences across proposals. Low scores in Market or Innovation reduce the final score more sharply than under a simple sum (MCTIC, 2018).

Figure 1. Flow of the three phases of the Centelha Program.
Figure 1. Flow of the three phases of the Centelha Program. Note. Prepared by the authors (2026), based on MCTIC (2018).
Table 1. Scoring criteria and formulas of the Centelha Program
PhaseCriteria and scaleScoring formula
Phase 1, Innovative IdeaMarket (M), Innovation (I) and Team (E); scale 0 to 6SCORE1 = (M × I) + E
Phase 2, Business ProjectInnovation (P) and Market (M): scale 4 to 10; Risk (R): 0.4 to 1.0SCORE2 = P × M × R
Phase 3, Funding ProjectProduct Planning (PP), Business Planning (PN), Team (E) and Budget (O); scale 4 to 10SCORE3 = (PP + PN + E + O) / 4
Final scoreComposition of phases 2 and 3FINAL SCORE = (SCORE2 + SCORE3) / 2

Note. Prepared by the authors based on MCTIC (2018).

Artificial intelligence and language models

Artificial intelligence is the field devoted to creating agents capable of perceiving the environment, reasoning and acting to achieve goals. Machine learning was consolidated by deep networks, while natural language processing was reconfigured by the Transformer architecture. On top of this architecture emerged large language models (LLMs), capable of executing linguistic tasks with few examples (Brown et al., 2020; Vaswani et al., 2017).

In the innovation funding domain, an LLM can analyze calls, extract criteria and structure technical texts by articulating dispersed information into coherent narratives. These capabilities coexist, however, with documented limitations such as factual hallucinations, difficulties with arithmetic and long-horizon planning, uneven use of long contexts and training biases. Table 2 summarizes these capabilities and limitations for proposal preparation (Bender et al., 2021; Ji et al., 2023; Liu, Lin, et al., 2024; Ooi et al., 2025; Valmeekam et al., 2023).

Table 2. Capabilities and limitations of LLMs for proposal preparation
DimensionCapabilitiesLimitations
Text generationCoherent and structured narratives; drafting of technical sectionsFactual hallucinations
Document analysisExtraction of criteria, requirements and deadlines from callsUneven effectiveness in long contexts
ReasoningProblem decomposition via Chain-of-ThoughtDifficulty with arithmetic and planning
AdaptationPerformance across diverse domains without retrainingBiases from training data

Multi-agent systems and cooperation patterns

An intelligent agent is an autonomous computational entity, situated in an environment, capable of perceiving and acting on it to achieve goals, with features of autonomy, sociability, reactivity and proactivity. A multi-agent system is the set of agents that coordinate actions and knowledge to solve problems that exceed individual capabilities. The distributed architecture provides robustness, scalability and fault tolerance compared with centralized systems (Dorri et al., 2018; Wooldridge, 2009).

Recent literature identifies three consolidated cooperation patterns in multi-agent systems, summarized in Table 3: voting-based, role-based and debate-based cooperation. Integration with LLMs has produced gains over single agents in textual revision and administrative automation. Recent research classifies agentic architectures and indicates that the most effective pattern combines hierarchical coordination, long-term memory and access to external tools, a configuration suited to tasks requiring multiple competencies such as the structuring of a funding proposal (Acharya et al., 2025; Liu, Lo, et al., 2025; Lu et al., 2024; Wu et al., 2024).

Table 3. Cooperation patterns in multi-agent systems
PatternMechanismAdvantagesTypical application
Voting-basedAgents vote to select solutionsFairness, equal distributionCollective decisions, alternative selection
Role-basedHierarchy with specialized rolesScalability, division of laborMultidisciplinary projects, technical drafting
Debate-basedArgumentation among agents for consensusAdaptability, transparencyText review, quality verification

Note. Synthesis based on Liu, Lo, et al. (2025) and Lu et al. (2024).

Mature applications of multi-agent systems with LLMs concentrate on healthcare, software engineering, recommender systems, supply chains and urban mobility. The application to the complete cycle of innovation funding proposal preparation remained unexplored before this work, a gap addressed by the product described below (Guo et al., 2024; He et al., 2025; Wu et al., 2024; Xi et al., 2023).

Methodology

The research is applied and technological in nature, exploratory in scope, and oriented by Design Science Research (DSR). DSR focuses on the creation and evaluation of artifacts to solve practical problems with a contribution to scientific knowledge, articulating methodological rigor and practical relevance. The operational path, illustrated in Figure 2, adopts the framework of Dresch et al. (2015a), composed of iterative stages that go from problem identification to communication of results, with feedback cycles between literature review, design and artifact evaluation (Dresch et al., 2015b; Hevner et al., 2004).

Figure 2. Elements of Design Science Research.
Figure 2. Elements of Design Science Research. Note. Adapted from Dresch et al. (2015).

The path was synthesized in four stages: the first, problem understanding, was based on analysis of the specialized literature on multi-agent systems, language models and public innovation funding, complemented by documentary review of the Centelha Program calls and regulations; the second involved artifact suggestion and development; the third comprised the empirical evaluation; and the fourth consisted of reflection and proposition formalization. The stages operate iteratively: findings from problem understanding fed refinements in the artifact during development, and the interviews produced incremental adjustments to agent prompts, initial information collection and the format of generated documents.

The artifact is classified as an instantiation, that is, a functional computational system that materializes concepts of multi-agent systems and generative AI. Development adopted Scrum in iterative two-week cycles. The initial core combined the Manager Agent and the WhatsApp webhook, expanded next to the coordinators and analysts. Implementation used Python and LangChain, connected to Google Gemini APIs. Orchestration was handled by the low-code n8n platform. Infrastructure runs on a cloud server in Docker containers, managed by EasyPanel. User communication occurs via WhatsApp Business API (Topsakal & Akinci, 2023).

The Scrum cycle was structured in biweekly sprints, each composed of planning, implementation, technical validation and review. The initial backlog prioritized integration with the WhatsApp Business API and the Manager Agent. Subsequent sprints progressively added the Phase 1, Phase 2 and Consolidation Coordinators, followed by Analysts specialized in market, technical, financial, legal and intellectual property dimensions. Each sprint closed with end-to-end functional tests over real Centelha call scenarios, recording prompt adjustments, orchestration flow refinements in n8n and MySQL persistence corrections. This pace allowed evolution from the minimum core to the full architecture of fifteen agents in approximately sixteen weeks.

The empirical evaluation was designed as an exploratory qualitative study and used semi-structured interviews with ten specialists (E1 to E10). Selection combined convenience and snowballing to ensure heterogeneous profiles: program evaluators and mentors, successful and rejected applicants, incubation managers, innovation consultants, regional support entity manager, university researchers and municipal public manager. Participants E1 and E2 composed the face validity test of the script; E3 to E10 formed the main sample. Interviews took place between February and April 2026 via Google Meet, after individual interaction with the artifact on WhatsApp (Creswell & Creswell, 2018). The sample size is consistent with exploratory qualitative designs oriented to analytical generalization, in which findings are generalized to theoretical propositions rather than to populations.

The script was organized in three blocks, totaling thirteen questions, whose complete wording is reproduced in Appendix A. Block A (questions Q1 to Q6) characterized the participants through categorical items on institutional affiliation, role and prior participation in the Centelha Program, and through three self-assessment items answered on five-point scales: involvement with the innovation ecosystem (Q3), familiarity with innovation funding calls (Q5) and familiarity with generative artificial intelligence tools (Q6). Block B (questions Q7 to Q12) evaluated the artifact against the official Centelha criteria: Innovative Idea (Q7), Business Project (Q8), Execution and Monitoring (Q9 and Q10) and Usability (Q11 and Q12). Questions Q11 and Q12 adapt items 3 and 1 of the System Usability Scale and were answered on five-point scales (Brooke, 1996). Block C consisted of an open synthesis question (Q13) on perceived contributions, limitations and suggestions. The open comments collected in questions Q7 to Q13 provided the corpus for the inductive coding described below, ensuring analytical depth and comparability across participants.

Interview analysis followed the Gioia, Corley and Hamilton method, with inductive category construction in three stages: first-order coding in informants’ language; grouping of codes into second-order theoretical themes; and synthesis of themes into aggregate dimensions. Triangulation was reinforced by the use of the consolidated data structure and the formalization of research propositions derived from emerging findings. Table 4 summarizes the methodological design (Gioia et al., 2013).

Table 4. Methodological design of the study
StageProcedureOutput
Problem understandingSystematic review in Scopus, Web of Science and Science Direct data-bases, with PRISMA protocolTwenty-nine selected papers; gap identified
Suggestion and developmentArchitecture modeling in five layers and fifteen agents; iterative Scrum implementationFunctional prototype of the multi-agent system
Empirical evaluationSemi-structured interviews with ten specialists after using the systemQualitative material for analysis
AnalysisGioia method (first- and second-order coding; aggregate dimensions)Four dimensions and four propositions

Classification rules for the consolidated results were defined prior to analysis. In Table 7, the Experience and AI familiarity columns report the participants’ self-assessments collected through the five-point scales of Block A of the script, reproduced in Appendix A, on the original scales. In Table 8, a criterion block was classified as full adherence when all ten specialists judged the corresponding outputs adequate in the open items, without substantive reservations; as adherence with reservations when at least one specialist conditioned adequacy on a specific correction, as occurred in the Business Project block with the review of budget figures (E2 and E4); and would have been classified as non-adherence if most specialists had judged the outputs inadequate, a situation not observed. These descriptive rules replace numerical cut-off points, which would presuppose a quantitative measurement model that this exploratory study does not claim.

Methodological rigor was reinforced by three safeguards. The first was triangulation between official program data, artifact execution records and perceptions of the ten specialists with complementary profiles. The second was face validity testing of the script with participants E1 and E2, whose feedback led to adjustments in question order and technical wording before the main round with E3 to E10. The third was systematic documentation of the Gioia coding process, with two review cycles of first- and second-order categories before aggregation into dimensions. Final propositions were derived exclusively from evidence consolidated in the data structure, in line with the interpretive validity criterion of qualitative research. Because the design is qualitative and exploratory, no quantitative measurement model was estimated and no model fit statistics apply; methodological quality rests on the safeguards described above and on the audit trail of the coding process.

Table 5. Functional layers of the multi-agent system architecture
LayerTechnologiesFunction
Interface and CommunicationWhatsApp Business API, n8n webhooks, WhisperMessage reception, media classification (text, audio, image, document), transcription
Orchestration and Flown8n (low-code)Sequential and conditional flows, agent activation and routing
Cognitive IntelligenceLangChain (AI Agent), Google Gemini 2.5 Flash LiteAgents with specialized prompts, own memory, text generation
Persistence and RetrievalMySQL (relational), PostgreSQL (vector)Structured project data and chat memory (100-message window)
Knowledge and RAGCall document base, embeddingsRetrieval-Augmented Generation for grounding agent responses

Note. Prepared by the authors (2026).

Table 6. Hierarchy of agents of the multi-agent system
LevelAgentFunction
1, StrategicManager AgentUser interface; progressive information collection; coordinator activation
2, TacticalPhase 1 CoordinatorManages the Innovative Idea analysis (M, I, E)
2, TacticalPhase 2 CoordinatorManages the Business Project (P, M, R)
2, TacticalPhase 3 CoordinatorManages the Funding Project (PP, PN, E, O)
2, TacticalConsolidation CoordinatorIntegrated review and final document preparation
3, OperationalMarket, Innovation and Team AnalystsPhase 1 criteria
3, OperationalFinancial, Product and Risk AnalystsPhase 2 criteria
3, OperationalBudget, IP, Compliance and Business AnalystsPhase 3 criteria
3, OperationalFinal Review AnalystSupport to the Consolidation Coordinator

Note. Prepared by the authors (2026).

Table 7. Characterization of empirical evaluation participants
IDProfileExperience (1 to 5)AI familiarity (1 to 5)
E1Centelha Program evaluator and mentor55
E2Successful Centelha applicant55
E3Innovation consultant43
E4Incubation program manager43
E5Early-stage entrepreneur22
E6Regional manager of entrepreneurship support entity22
E7Applicant with previously rejected proposal54
E8University researcher in technology transfer33
E9Consultant specialized in tax-incentive laws54
E10Municipal director of Science, Technology and Innovation42

Note. Prepared by the authors (2026).

Table 8. Summary of findings by Centelha criterion
CriterionResultEvidence (interviews)
Innovative IdeaFull adherence in 10 of 10 participantsE8 highlighted that the system brought many new aspects to the theme. E9 stated that it develops, unfolds and deepens in additional layers
Business ProjectAdherence with reservations in 2 of 10 (financial figures)E2 and E4 pointed to the need for human review of budget components, in line with arithmetic limitations of LLMs
Execution and MonitoringAdherence to criteria; coherent structure across phasesE7, with a previously rejected proposal, reported that with little infor-mation the system produced a more complex and explanatory unfolding
UsabilityUnanimous positive evaluation; recurrent use intentE6 projected demand for fifteen to eighteen new projects per week at the regional office. E10 considered the interface adaptable to the insti-tutional context

Note. Prepared by the authors (2026).

Interviews were recorded with explicit participant consent, transcribed and coded with the support of structured spreadsheets. Coding followed an internal pairwise review routine among the authors, with discussion of divergences until consensus. A theoretical saturation criterion was adopted to close the main round: from the eighth interview (E8) onwards, newly emerging categories no longer expanded the consolidated data structure, indicating stabilization of the aggregate dimensions. The last two specialists (E9 and E10) confirmed this pattern and reinforced, respectively, the dimensions of institutional transferability and territorial public governance. The research observed the ethical principles of human research, ensuring participant anonymity, voluntariness, possibility of withdrawal at any time and restricted custody of recordings and transcripts.

Results

Architecture and implementation of the artifact

The technical-technological product described in this paper falls within the Software category, in line with the classification adopted by the Brazilian Coordination for the Improvement of Higher Education Personnel for Professional Master’s programs. It is an operational prototype of a multi-agent system based on generative AI to support the preparation of innovation funding projects, focused on the Centelha Program. The artifact is documented on a public hotsite that provides a usage tutorial, architecture description and a direct interaction channel with the agent (Author, 2026).

The architecture was conceived in five functional layers, an organization that ensures modularity, scalability and traceability and aligns with reference architectures for foundation-model agents and the containerization pattern for open multi-agent systems. Table 5 details the layers, their technologies and functions (Lima & Aguiar, 2024; Liu, Lo, et al., 2025; Lu et al., 2024).

The cognitive core comprises fifteen agents distributed across three hierarchical levels, in role-based cooperation. This choice aligns with the nature of the problem, which demands division of labor among distinct specialties such as market, technical, financial, legal and intellectual property analysis. Each agent is implemented as a cognitive microservice through the LangChain AI Agent tool inside n8n, with independent chat memory, a dedicated MySQL persistence tool and a specialized prompt. The Google Gemini 2.5 Flash Lite model was selected as the cognitive engine to balance operational cost and generation quality.

The implementation strategy evolved during development. The original plan envisaged independent Python microservices; the final version adopted native LangChain/n8n integration, which favored simplicity and rapid iteration without compromising modularity. Table 6 details the agent hierarchy; Figure 3 presents the consolidated view of the architecture, integrating layers and agents (Topsakal & Akinci, 2023).

Figure 3. Multi-agent system architecture: functional layers and agent hierarchy.
Figure 3. Multi-agent system architecture: functional layers and agent hierarchy. Note. Prepared by the authors (2026).

The cycle begins when the user sends a message via WhatsApp. The n8n webhook receives the input and applies idempotency filters and exclusion of system-generated messages. The input goes through media classification and, when applicable, Whisper transcription for audio or text extraction for PDFs and images. The Manager Agent conducts the collection of information about the business idea and, once enough data is accumulated, sequentially triggers the Phase Coordinators. Each coordinator delegates to the competent analysts the structuring of textual blocks corresponding to the official criteria. The Consolidation Coordinator performs the final integration, and the document is sent by email. Retrieval-Augmented Generation techniques allow agents to consult call documents during generation, reducing factual hallucinations (Ji et al., 2023).

The artifact, in functional state, materializes the concept of a team of specialists in which each agent assumes a specific role, articulated by an orchestrator agent that maintains contextual continuity throughout the process. In addition to the software, a hotsite was produced describing the program, the architecture and the research methodology, and offering a direct interaction channel with the system.

Empirical evaluation

This subsection presents the results of the exploratory evaluation of the artifact with ten specialists, conducted between February and April 2026. The professional profiles are heterogeneous and cover different positions in the funding cycle, as detailed in Table 7. Analysis is organized by the official Centelha criteria, grouped in four blocks: Innovative Idea, Business Project, Execution and Monitoring, and Usability.

Table 8 summarizes the consolidated findings by criterion. Adherence to the Innovative Idea block was full across the ten participants. E8 highlighted that the system brought many new aspects to the theme, and E9 stated that it develops, unfolds and deepens in additional layers. The Business Project block showed adherence with reservations in 2 of 10 interviewees (E2 and E4), who pointed to the need for human review of budget components, in line with the arithmetic limitations of LLMs. The Execution and Monitoring block showed a coherent structure across phases; E7, with a previously rejected proposal, reported that, with little information, the system produced a more complex and explanatory unfolding. The Usability block received unanimous positive evaluation and indication of recurrent use; E6 projected demand for fifteen to eighteen new projects per week at the regional office, and E10 considered the interface adaptable to the institutional context.

Inductive analysis consolidated four aggregate dimensions, summarized below and detailed in Table 9 (Gioia et al., 2013). The first dimension, Perceived value of the solution, groups perceptions on channel accessibility, artifact maturity and scalability potential. The second dimension, Quality of the generated outputs, evidences the coherence and completeness of the produced documents, the occasional occurrence of hallucinations in quantitative values and the cross-cutting theme of responsible use and human review, reinforced by the report of E10, who described concrete cases of AI authorship detection in state calls. The third dimension, User experience, captures perceptions on initial guidance, responsiveness and context adaptation. The fourth dimension, Extension potential, summarizes educational and social applicability, technical-operational feasibility and transferability to other programs and institutions.

Table 9. Consolidated data structure from the Gioia method
Aggregate dimensionSecond-order themesFirst-order codes (examples)
1. Perceived value of the solutionChannel accessibility; product maturity; scalabilityUnfolds and deepens in layers (E9); WhatsApp familiarity fa-cilitates (E5, E10)
2. Quality of the generated outputsTextual coherence; hallucinations in figures; responsible use and human reviewMissing value check (E2, E4); we need to review before sub-mission (E10)
3. User experienceGuidance and onboarding; responsiveness; context adap-tationQuite investigative (E10); adapts to context (E10); natural flow (E5, E6)
4. Extension potentialEducational application; institutional feasibility; transfe-rabilityWorks for Lei do Bem (E9); fifteen to eighteen projects per week (E6); fundraising intelligence (E10)

Note. Prepared by the authors (2026), based on Gioia et al. (2013).

From the four dimensions, four theoretical propositions emerged, formalized for testing in future comparative quantitative studies. Proposition P1 holds that the use of MAS with LLMs reduces the asymmetry of technical-argumentative capital among applicants, levelling articulation capacities in international and Brazilian programs. Proposition P2 holds that ubiquitous conversational interfaces, such as WhatsApp, expand adoption intent compared with traditional web interfaces, particularly in territories with digital infrastructure constraints. Proposition P3, stated as a hypothesis for future comparative research, conjectures that the hierarchical multi-agent architecture can be transferred to other funding programs, Brazilian (Lei Rouanet, Lei do Bem, Funcad, FAPESP PIPE) and international (SBIR Phase I, EIC Accelerator, Innovate UK Smart Grants), with predominantly parametric adaptations; its empirical basis is restricted to the assessments of ten specialists within a single program, and transferability therefore remains to be demonstrated. Proposition P4, likewise a hypothesis rather than a demonstrated claim, conjectures that the integration of MAS with generative AI in municipal and state instances may reconfigure the role of local public authorities as an infrastructure of intelligence for territorial-scale resource attraction, a possibility raised by the public-sector participants (E6, E10) that requires multi-site evidence before any generalization. The propositions converge with findings on hierarchical architectures for tasks requiring coordination across specialties (Acharya et al., 2025; Lima & Aguiar, 2024; Liu, Lo, et al., 2025; Lu et al., 2024).

Discussion

The proposed architecture contributes on three fronts. On the theoretical front, the research fills a gap identified in the literature by integrating multi-agent systems and LLMs into the complete cycle of innovation funding proposal preparation, a domain so far not addressed by the combined literature. The three-level hierarchical organization materializes role-based cooperation. Context adaptation articulates market, technical, financial, legal and intellectual property analysis. The formalization of the four propositions opens an investigative front for comparative quantitative studies and expands the theoretical base of AI applied to innovation management (Gama & Magistretti, 2025; Guo et al., 2024; Liu, Lo, et al., 2025; Lu et al., 2024; Wang et al., 2024; Xi et al., 2023).

On the practical front, the system has the potential to reduce the time and effort devoted to fundraising. Modularity and the use of open standards, such as Docker, REST APIs and LangChain, favor extensibility to Brazilian and international funding programs, as indicated by participants E9 (consultant) and E10 (public manager). Operation at marginal cost enables business models such as public SaaS or a tool embedded in funding agency platforms (Gama & Magistretti, 2025; Mazzucato, 2018). An embedded deployment of this kind would extend the articulation between funding agencies, firms and research institutions documented in the recent generations of Brazilian programs (Bezerra Borges et al., 2021) and would attack the access barriers that keep the effective uptake of Brazilian instruments below their potential (Fabiani & Sbragia, 2014).

On the social front, the artifact levels technical articulation capacities among applicants with different degrees of familiarity with the language of calls. There is potential to democratize access to funding programs and broaden the diversity of ideas and territories represented in public innovation funding systems, in line with the mission orientation of the recent innovation policy literature (Edler & Fagerberg, 2017; Mazzucato, 2018). The statement of participant E7, whose proposal had been rejected in a previous edition, synthesizes this reach: with little information, the system produced a more complex and explanatory unfolding of the project. This testimony empirically operationalizes Proposition P1.

The participation of the ten specialists constituted a decisive factor in qualifying the artifact. The complementary profiles covered the funding cycle in different positions, such as evaluation, mentoring, application, incubation management, consultancy and public governance. Each interview brought concrete adjustments to agent prompts, information collection and output format, contributing to product maturation. The analytical depth of the Gioia method transformed this diversity into four formalized theoretical propositions, expanding the scientific reach of the work.

The four derived propositions establish a research agenda that dialogues with the recent literature on foundation-model agents. Proposition P1, on the reduction of asymmetries among applicants, provides a basis for quantitative studies comparing approval rates before and after artifact use. Proposition P2, on the ubiquitous channel, connects with the literature on technology acceptance in territories with low digital infrastructure and can be operationalized through the application of the System Usability Scale in diverse cohorts. Proposition P3, on transferability, is consistent with the typology of AI applications in innovation management and suggests that the agent design pattern catalogue consolidated by the recent literature may be reapplicable to other calls, a conjecture that comparative studies across programs will need to test (Gama & Magistretti, 2025; Liu, Lo, et al., 2025; Wang et al., 2024). Proposition P4, on the possible reconfiguration of the role of municipal and state instances, adds, as a hypothesis grounded in the perceptions of the public-sector participants, to the debate on hierarchical multi-agent architectures as public infrastructures of applied intelligence (Guo et al., 2024; Lu et al., 2024; Xi et al., 2023). In the Brazilian context, where Industry 4.0 research advances more slowly than in leading countries and emphasizes operational efficiency (Scala da Rocha & De Oliveira Matias, 2026), artifacts of this kind also connect innovation management research with the national digital transformation agenda.

Compared with other multi-agent systems with LLMs, the proposed artifact differs by combining a ubiquitous conversational channel, agent-segmented persistence and a low-code orchestration flow. Platforms oriented to inter-agent conversation for programming or administrative support tasks prioritize internal autonomy and machine-to-machine iteration, while the prototype described here keeps the human at the center of decision-making and adopts a messaging application as the entry point (Wu et al., 2024). This choice aligns with the responsibility principle of the reference architecture for foundation-model agents, under which operational transparency and the possibility of user review are central requirements for deployment in sensitive domains (Lu et al., 2024).

An ethical tension runs through these findings and deserves explicit treatment: the artifact assists the generation of funding proposals with generative AI, while participant E10, in the same study, reported concrete cases of proposals disqualified in state-level programs after the detection of AI authorship. The two observations can only be reconciled under an explicit governance protocol, which the findings of this study allow us to outline in three commitments. The first is a mandatory human-review protocol: the system’s output is a structured draft, and the proposing team must review, correct and complement every section before submission — in particular the quantitative content, such as budgets and schedules, for which hallucinations were observed (E2, E4) — so that final factual and argumentative responsibility remains with the applicants. The second is a threshold of editorial intervention over the artifact’s output: the system is designed to organize information supplied by the applicant during the conversational flow, not to invent content on the applicant’s behalf; substantive claims about market, technology and team must originate from the user and be verifiable, and the generated text is to be treated as an editable scaffold in which the applicant’s own voice replaces machine phrasing wherever judgement is involved. The third is transparency vis-à-vis the evaluating agency: use of the system should be disclosed whenever the call requires it or an AI-use policy exists, and, where a program explicitly prohibits AI-generated text, the artifact’s role must retreat to call analysis, criteria explanation and structural guidance, leaving drafting entirely to the applicant. Under these conditions, using the system becomes analogous to hiring specialized grant-writing support, a long-accepted practice, whereas full delegation without review remains both an ethical violation and, as E10’s report shows, a concrete risk of disqualification. This protocol operationalizes the responsibility principle of reference architectures for foundation-model agents (Lu et al., 2024) and converts the tension documented in the findings into an explicit condition of use of the artifact.

Five limitations stand out, the first two interlinked. The first is the occasional occurrence of hallucinations in LLM outputs, particularly in quantitative values such as financial estimates and technology readiness levels. The second is the need for human validation of generated proposals before submission, given the observed tendency of applicants to fully delegate writing to artificial intelligence; the governance protocol outlined above is the direct answer to this risk. The third limitation concerns the sample: the exploratory design with ten specialists in a single program supports analytical generalization to the propositions formalized in Section 4 and does not support statistical generalization to the population of applicants. The fourth is the perceptual nature of the evaluation: the evidence consists of specialists’ judgements of the artifact’s outputs against the official criteria, not of causal measures of effect; in particular, no proposal generated with the system was actually submitted to an official Centelha selection round, so effects on approval rates remain unmeasured. The fifth is the dependence on a proprietary commercial foundation model (Google Gemini 2.5 Flash Lite), with implications for cost, data sovereignty, reproducibility and service continuity, only partially mitigated by the portability of the architecture across providers. Mitigation of the first two limitations depends on practices such as Retrieval-Augmented Generation, automatic numerical consistency checks and explicit human review protocols before submission; the third and fourth limitations define the expanded validation planned in Section 6. Table 10 summarizes contributions and limitations by dimension (Bender et al., 2021; Ji et al., 2023; Valmeekam et al., 2023).

Table 10. Contributions and limitations of the technical-technological product
DimensionContributionLimitation
TheoreticalMAS architecture with LLM for funding; four formalized propositions (P1 to P4)Need for systematic human validation of outputs before submission
TechnicalFifteen agents in three hierarchical levels; open standardOccasional LLM hallucinations, especially in quantitative values
EconomicMarginal operational cost enables scaleDependence on commercial providers of foundation models
SocialReduction of asymmetries among applicants; democratization of fundingEthical risk of fully delegating writing to AI without human review
UsabilityConversational interface via WhatsAppChannel limitations for complex interactions and long flows

Note. Prepared by the authors (2026).

Conclusions

This research presented the development and exploratory empirical evaluation of a technical-technological product in the Software category: a multi-agent system based on generative AI supporting proposal preparation for public innovation funding programs, with application to the Centelha Program as the Brazilian case. The artifact was designed under Design Science Research and implemented in an architecture of five functional layers and fifteen hierarchical agents. The evaluation involved ten specialists with complementary profiles, and inductive analysis using the Gioia method consolidated four aggregate dimensions and formalized four research propositions.

Contributions are articulated on three planes. Theoretically, the work fills a specific gap at the intersection of multi-agent systems, large language models and public innovation funding, proposing a three-level hierarchical architecture and formalizing propositions P1 through P4. Practically, the integration of call analysis, structured textual generation and a ubiquitous conversational interface enables reductions in time and effort devoted to fundraising, at marginal operational cost. Socially, the testimony of E7, whose submission had been rejected, illustrates the reach of the artifact as a reducer of asymmetries of technical-argumentative capital among applicants with different degrees of familiarity with the language of calls.

Implications for innovation policy extend beyond the Brazilian case. The architecture is designed from elements common to early-stage subvention programs in different jurisdictions, such as SBIR/STTR, the EIC Accelerator, Innovate UK Smart Grants and the Centelha Program. Transferability indicated by specialists E9 and E10 suggests that funding agencies and municipal or state managers could operate adapted instances as a public infrastructure of intelligence applied to fundraising. The potential impact, conditional on future testing of the propositions, is threefold: greater diversity of applicants and territories in the public funding system, better average quality of submitted proposals and efficiency gains for official evaluators.

Limitations stand out on five fronts. The first is the occasional occurrence of hallucinations in LLM outputs, particularly in quantitative values. The second is the need for systematic human validation of generated proposals before submission, under the governance protocol discussed in Section 5. The third is the current dependence on a proprietary commercial foundation model, with implications for cost, data sovereignty and service continuity. The fourth is the exploratory scope of the evaluation, based on the perceptions of ten specialists in a single program, which restricts the findings to analytical generalization and keeps propositions P1 to P4 as hypotheses for future quantitative testing. The fifth is that the evaluation did not include the actual submission of generated proposals to an official selection round, so effects on approval rates remain unmeasured. Mitigation depends on practices such as Retrieval-Augmented Generation, automatic numerical consistency checks and the human review protocol before submission.

The future agenda comprises three fronts. The first is agent self-configuration, in which the system reads a call and automatically models the number and specialization of agents, making the architecture adaptive to international programs. The second is expanded validation, with submission of proposals to official evaluators in a blind scenario and application of the System Usability Scale to a larger sample, enabling a quantitative measurement model with reported fit statistics and validated cut-off points. The third is expansion to other Brazilian programs (FAPESP PIPE, Lei Rouanet, Funcad, Lei do Bem) and international ones (SBIR Phase I, EIC Accelerator, Innovate UK Smart Grants).

Appendix A. Semi-structured interview script

The script below was applied between February and April 2026, after each specialist interacted individually with the artifact on WhatsApp. Participants E1 and E2 composed the face validity test; E3 to E10 formed the main round.

Block A. Participant characterization

Q1. Type of institution in which the participant works (government; university or science and technology institution; company; other).

Q2. Level of participation in the Centelha Program (never participated; submitted proposals; acted as evaluator; program management; other).

Q3. On a scale from 1 to 5, degree of involvement with the innovation ecosystem (1 = very low; 5 = very high).

Q4. Role in the organization (technical staff; manager; consultant; researcher; other).

Q5. On a scale from 1 to 5, familiarity with innovation funding calls (1 = has never read a call; 5 = analyzes calls routinely).

Q6. On a scale from 1 to 5, familiarity with generative artificial intelligence tools (1 = has never used them; 5 = uses them daily).

Block B. Artifact evaluation by official Centelha criterion

Q7. Innovative Idea criterion (Phase 1). After the interaction with the system, does the generated document describe clearly and convincingly the innovative and differentiating elements of the proposed solution? Comment on strengths and limitations.

Q8. Business Project criterion (Phase 2). Does the generated document adequately structure the technical and financial execution plan (product, business, team, budget)? Identify inconsistencies, if any.

Q9. Execution and Monitoring criterion (Phase 3), concrete results. Does the document present a schedule, a risk mitigation plan and a feasibility analysis compatible with the nature of the project?

Q10. Execution and Monitoring criterion (Phase 3), adjustments and justifications. Does the document make it possible to justify implementation adjustments during project execution? Does the model show adequate flexibility?

Q11. Usability criterion, ease of use (System Usability Scale item 3). Is the system easy to use, without the need for external assistance? Answered on a scale from 1 to 5.

Q12. Usability criterion, frequency of use (System Usability Scale item 1). Would you use this system frequently to support new submissions? Answered on a scale from 1 to 5.

Block C. Synthesis of the experience

Q13. Considering the complete interaction with the system, describe freely: (i) the main perceived contributions; (ii) the limitations identified; (iii) suggestions for improvement and future applications of the artifact.

Note. Answers to questions Q7 to Q12 combined the objective scales with open comments, which constituted the corpus for the coding by the Gioia method described in Section 3.

References

  1. Acharya, D. B., Kuppan, K., & Divya, B. (2025). Agentic AI: Autonomous intelligence for complex goals—A comprehensive survey. IEEE Access, 13, 18912–18936. https://doi.org/10.1109/ACCESS.2025.3532853
  2. Arrow, K. J. (1962). Economic welfare and the allocation of resources for invention. In R. R. Nelson (Ed.), The rate and direction of inventive activity: Economic and social factors (pp. 609–626). Princeton University Press. https://doi.org/10.1515/9781400879762-024
  3. Author. (2026). Multi-agent system based on generative AI to support innovation funding projects: Institutional hotsite [Reference redacted for blind review].
  4. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. https://doi.org/10.1145/3442188.3445922
  5. Bezerra Borges, D., Meyer Soares, P., & Santana Silva, M. (2021). Programs and instruments for promoting innovation with technology-based companies in Brazil. Journal of Technology Management & Innovation, 16(2), 28–40. https://doi.org/10.4067/S0718-27242021000200028
  6. Brooke, J. (1996). SUS: A “quick and dirty” usability scale. In P. W. Jordan, B. Thomas, B. A. Weerdmeester, & I. L. McClelland (Eds.), Usability evaluation in industry (pp. 189–194). Taylor & Francis.
  7. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901. https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html
  8. Castro, G. F. O., Rosário, F. J. P., & Lima, A. A. (2022). O Programa Centelha AL como fonte de inovação frugal [The Centelha AL Program as a source of frugal innovation]. Diversitas Journal, 7(1), 376–389. https://doi.org/10.48017/dj.v7i1.2066
  9. Creswell, J. W., & Creswell, J. D. (2018). Research design: Qualitative, quantitative, and mixed methods approaches (5th ed.). Sage.
  10. Dorri, A., Kanhere, S. S., & Jurdak, R. (2018). Multi-agent systems: A survey. IEEE Access, 6, 28573–28593. https://doi.org/10.1109/ACCESS.2018.2831228
  11. Dresch, A., Lacerda, D. P., & Antunes Júnior, J. A. V. (2015a). Design science research: Método de pesquisa para avanço da ciência e tecnologia [Design science research: A research method for the advancement of science and technology]. Bookman.
  12. Dresch, A., Lacerda, D. P., & Miguel, P. A. C. (2015b). Uma análise distintiva entre o estudo de caso, a pesquisa-ação e a design science research [A distinctive analysis of case study, action research and design science research]. Revista Brasileira de Gestão de Negócios, 17(56), 1116–1133. https://doi.org/10.7819/rbgn.v17i56.2069
  13. Edler, J., & Fagerberg, J. (2017). Innovation policy: What, why, and how. Oxford Review of Economic Policy, 33(1), 2–23. https://doi.org/10.1093/oxrep/grx001
  14. Fabiani, S., & Sbragia, R. (2014). Tax incentives for technological business innovation in Brazil: The use of the Good Law – Lei do Bem (Law No. 11196/2005). Journal of Technology Management & Innovation, 9(4), 53–63. https://doi.org/10.4067/S0718-27242014000400004
  15. Gama, F., & Magistretti, S. (2025). Artificial intelligence in innovation management: A review of innovation capabilities and a taxonomy of AI applications. Journal of Product Innovation Management, 42(1), 76–111. https://doi.org/10.1111/jpim.12698
  16. Gioia, D. A., Corley, K. G., & Hamilton, A. L. (2013). Seeking qualitative rigor in inductive research: Notes on the Gioia methodology. Organizational Research Methods, 16(1), 15–31. https://doi.org/10.1177/1094428112452151
  17. Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., & Zhang, X. (2024). Large language model based multi-agents: A survey of progress and challenges. arXiv. https://doi.org/10.48550/arXiv.2402.01680
  18. He, J., Treude, C., & Lo, D. (2025). LLM-based multi-agent systems for software engineering: Literature review, vision, and the road ahead. ACM Transactions on Software Engineering and Methodology, 34(5), Article 124. https://doi.org/10.1145/3712003
  19. Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75–105. https://doi.org/10.2307/25148625
  20. Howell, S. T. (2017). Financing innovation: Evidence from R&D grants. American Economic Review, 107(4), 1136–1164. https://doi.org/10.1257/aer.20150808
  21. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730
  22. Lerner, J. (2010). The future of public efforts to boost entrepreneurship and venture capital. Small Business Economics, 35(3), 255–264. https://doi.org/10.1007/s11187-010-9298-z
  23. Lima, G. L., & Aguiar, M. S. (2024). Towards a Docker-based architecture for open multi-agent systems. IAES International Journal of Artificial Intelligence, 13(1), 45–56. https://doi.org/10.11591/ijai.v13.i1.pp45-56
  24. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638
  25. Liu, Y., Lo, S. K., Lu, Q., Zhu, L., Zhao, D., Xu, X., Harrer, S., & Whittle, J. (2025). Agent design pattern catalogue: A collection of architectural patterns for foundation model based agents. Journal of Systems and Software, 220, Article 112278. https://doi.org/10.1016/j.jss.2024.112278
  26. Lu, Q., Zhu, L., Xu, X., Xing, Z., Harrer, S., & Whittle, J. (2024). Towards responsible generative AI: A reference architecture for designing foundation model based agents. 2024 IEEE 21st International Conference on Software Architecture Companion (ICSA-C), 119–126. https://doi.org/10.1109/ICSA-C63560.2024.00028
  27. Mazzucato, M. (2018). Mission-oriented innovation policies: Challenges and opportunities. Industrial and Corporate Change, 27(5), 803–815. https://doi.org/10.1093/icc/dty034
  28. Ministério da Ciência, Tecnologia, Inovações e Comunicações. (2018). Portaria n° 4.082, de 10 de agosto de 2018: Institui o Programa Nacional de Apoio à Geração de Empreendimentos Inovadores (Programa Centelha) [Ordinance No. 4,082 of August 10, 2018: Establishing the National Program to Support the Generation of Innovative Enterprises (Centelha Program)]. https://antigo.mctic.gov.br/mctic/opencms/legislacao/portarias/Portaria_MCTIC_n_4082_de_10082018.html
  29. Ooi, K.-B., Tan, G. W.-H., Al-Emran, M., Al-Sharafi, M. A., Capatina,A., Chakraborty, A., Dwivedi, Y. K., Huang, T.-L., Kar, A. K., Lee, V.-H., Loh, X.-M., Micu, A., Mikalef, P., Mogaji, E., Pandey, N., Raman,R., Rana, N. P., Sarker, P., Sharma, A., … Wong, L.-W. (2025). The potential of generative artificial intelligence across disciplines: Perspectives and future directions. Journal of Computer Information Systems, 65(1), 76–107. https://doi.org/10.1080/08874417.2023.2261010
  30. Rapini, M. S., & Rocha, B. P. (2014). Bancos de desenvolvimento e o financiamento da inovação [Development banks and the financing of innovation]. Caderno Econômico BDMG, (2), 7–58. https://www. bdmg.mg.gov.br/wp-content/uploads/2018/10/BDMG-Caderno-Econ%C3%B4mico-Dezembro-2014-N%C2%B0-2.pdf
  31. Scala da Rocha, L., & De Oliveira Matias, Í. (2026). Industry 4.0 research in emerging economies: A bibliometric comparison of Brazil and globally leading countries. Journal of Technology Management & Innovation, 21(1), 112–125. https://doi.org/10.4067/S0718-27242026000100112
  32. Schumpeter, J. A. (1934). The theory of economic development: An inquiry into profits, capital, credit, interest, and the business cycle (R. Opie, Trans.). Harvard University Press.
  33. Topsakal, O., & Akinci, T. C. (2023). Creating large language model applications utilizing LangChain: A primer on developing LLM apps fast. Proceedings of the 5th International Conference on Applied Engineering and Natural Sciences (ICAENS), 1050–1056. https://doi.org/10.59287/icaens.1127
  34. Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., & Kambhampati, S. (2023). PlanBench: An extensible benchmark for evaluating large language models on planning and reasoning about change. Advances in Neural Information Processing Systems, 36, 38975–38987. https://proceedings.neurips.cc/paper_files/paper/2023/hash/7a92bcdede88c7afd108072faf5485c8-Abstract-Datasets_and_Benchmarks.html
  35. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez,A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
  36. Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z.,Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., & Wen, J. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6), Article 186345. https://doi.org/10.1007/s11704-024-40231-1
  37. Wooldridge, M. (2009). An introduction to multiagent systems (2nd ed.). John Wiley & Sons.
  38. Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X.,Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., & Wang,C. (2024). AutoGen: Enabling next-gen LLM applications via multiagent conversations. Proceedings of the 1st Conference on Language Modeling (COLM). https://openreview.net/forum?id=BAakY1hNKS
  39. Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., … Gui, T. (2023). The rise and potential of large language model based agents: A survey. arXiv. https://doi.org/10.48550/arXiv.2309.07864