<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="part2stratml.xsl"?>
<PerformancePlanOrReport xmlns="urn:ISO:std:iso:17469:tech:xsd:PerformancePlanOrReport" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"

 xsi:schemaLocation="urn:ISO:std:iso:17469:tech:xsd:PerformancePlanOrReport http://stratml.us/references/PerformancePlanOrReport20160216.xsd" Type="Strategic_Plan"><Name>NAIRR Structure and Specifications for Resource Elements</Name><Description>This chapter provides details of [the] key components, along with desired capabilities when the NAIRR begins initial operations. Given the fast pace of technological development, the Operating Entity should maintain the flexibility to adjust approaches to the elements detailed below, in consultation with the Steering Committee and Program Management Office.</Description><OtherInformation>The NAIRR Operating Entity should develop an integrated portal to provide the user base described in Chapter 2 with access to a federated mix of on-premise and commercial computational and data resources and services. Computational resources would include conventional servers, computing clusters, HPC, and cloud computing, and should also support access to edge computing resources and testbeds for AI R&amp;D. The NAIRR Operating Entity should make open and protected data available via resource providers and partnerships. Data should be co-located with computational resources where possible. Data providers should facilitate user access to restricted statistical data through the Standard Application Process (SAP) established under the 2018 Foundations for Evidence-Based Policymaking Act, where appropriate and possible.
^^
The NAIRR Operating Entity and resource providers should make software, training, and educational resources available to support a diverse set of users with varying levels of AI research experience and proficiency. </OtherInformation><StrategicPlanCore><Organization><Name>National Artificial Intelligence Research Resource Task Force</Name><Acronym>NAIRRTF</Acronym><Identifier>_6ca63702-ceaf-11ed-b4fb-9e69fe82ea00</Identifier><Description/><Stakeholder StakeholderTypeType="Person"><Name/><Description/></Stakeholder></Organization><Vision><Description>A widely-accessible, national cyberinfrastructure that will advance and accelerate the U.S. AI R&amp;D environment and fuel AI discovery and innovation in the United States</Description><Identifier>_6ca638a6-ceaf-11ed-b4fb-9e69fe82ea00</Identifier></Vision><Mission><Description>To provide details of the key components of the NAIRR along with desired capabilities</Description><Identifier>_48f034ca-cf63-11ed-b3e1-ed4b1883ea00</Identifier></Mission><Value><Name>Artificial Intelligence</Name><Description/></Value><Value><Name>Research</Name><Description/></Value><Value><Name>Development</Name><Description/></Value><Value><Name>Capability</Name><Description/></Value><Goal><Name>Portal &amp; Interface</Name><Description>Develop an NAIRR user portal</Description><Identifier>_48f03696-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>Access Portal and User Interface ~ The Operating Entity is responsible for development of an NAIRR user portal that supports key user functionalities such as single sign-on, team allocations, data search and discovery, collaboration tools, resource discovery, job submission, consolidated accounting, spend alerts, information about data use, and cost-optimization of workflows. The portal will be one way to access NAIRR resources. Alternate access methods such as secure shell or scripting interfaces should also be made available for advanced users.</OtherInformation><Objective><Name>Catalog &amp; Jobs</Name><Description>Allow users to select AI applications, computational resources, and data sources from a curated catalog and to launch and monitor jobs</Description><Identifier>_48f037c2-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The portal will allow users to select their AI applications, computational resources, and data sources from a curated catalog, and to launch and monitor jobs from a portal that provides a uniform, integrated view.</OtherInformation></Objective><Objective><Name>Help</Name><Description>Provide built-in help functions and an integrated help desk ticketing system</Description><Identifier>_48f038d0-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.2</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The portal should have built-in help functions and an integrated help desk ticketing system.</OtherInformation></Objective><Objective><Name>Documentation &amp; Training</Name><Description>Maintain a catalog of resource provider user documentation and training materials</Description><Identifier>_48f039de-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.3</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The portal should maintain an up-to-date catalog of resource provider user documentation and training materials.</OtherInformation></Objective><Objective><Name>Collaboration &amp; Communities</Name><Description>Support collaboration and community building</Description><Identifier>_48f03baa-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Students</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Researchers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Resource Providers</Name><Description/></Stakeholder><OtherInformation>Chat functions, meeting rooms, forums, and other functionality may be included to support collaboration and community building among students, researchers, resource providers, and other users.</OtherInformation></Objective><Objective><Name>Search &amp; Discovery</Name><Description>Enable data search and discovery and leverage automated technologies</Description><Identifier>_48f03cc2-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.5</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The portal should also enable data search and discovery and leverage automated technologies so that (1) metrics on data use can drive data acquisition and (2) diverse, community-driven data curation, linkage, and validation activities can be fostered.</OtherInformation></Objective><Objective><Name>Metrics</Name><Description>Use metrics on data use to drive data acquisition</Description><Identifier>_d978c784-d29a-11ed-adce-039f0c83ea00</Identifier><SequenceIndicator>1.5.1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Curation, Linkage &amp; Validation</Name><Description>Foster community-driven data curation, linkage, and validation activities</Description><Identifier>_d978c928-d29a-11ed-adce-039f0c83ea00</Identifier><SequenceIndicator>1.5.2</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>User Accounts</Name><Description>Maintain user accounts</Description><Identifier>_48f03dd0-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.6</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>A user account would be required to manage computational allocations, monitor usage, submit jobs, and post to the community forum.</OtherInformation></Objective><Objective><Name>Website</Name><Description>Provide a website through which key elements are publicly available</Description><Identifier>_48f03ed4-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.7</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The Operating Entity should provide a public website through which some key elements are available without the need for a user account and sign-on. For example, linked catalogs of AI education tools and testbeds, as well as an index of AI datasets with metadata, annotations of known problems and deprecation status, and community-contributed code, should be readily available.</OtherInformation></Objective><Objective><Name>Cost</Name><Description>Assess the cost of building the user portal and public website inhouse versus acquiring it commercially</Description><Identifier>_48f03ff6-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>1.8</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The Operating Entity should assess the cost of building the user portal and public website inhouse versus acquiring it commercially. To speed development, the Operating Entity could outsource the design, construction, and maintenance of the user portal to a commercial entity that has previously created successful user portals. All major aspects of the portal should be included in NAIRR initial operational capabilities.</OtherInformation></Objective></Goal><Goal><Name>Access</Name><Description>Make access to computational and data resources available to a variety of new users who otherwise would face financial, logistical, or capacity challenges engaging in the AI research ecosystem</Description><Identifier>_48f04104-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>2</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>Computational Resources ~ To lower barriers to entry into AI research, the Operating Entity and resource providers must make access to computational and data resources available to a variety of new users who otherwise would face financial, logistical, or capacity challenges engaging in the AI research ecosystem.</OtherInformation><Objective><Name>Leverage</Name><Description>Leverage existing resources in all sectors</Description><Identifier>_48f04212-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>2.1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>Expanded access should be provided by leveraging existing resources in all sectors, augmenting the capacity of federally provided resources as appropriate, creating new research computing and data infrastructure to serve the AI R&amp;D community, and providing financial support where needed.</OtherInformation></Objective><Objective><Name>Federal Resources</Name><Description>Augment the capacity of federally provided resources as appropriate</Description><Identifier>_48f0437a-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>2.1.1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Infrastructure</Name><Description>Create new research computing and data infrastructure to serve the AI R&amp;D community</Description><Identifier>_48f04492-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>2.1.2</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Financial Support</Name><Description>Provide financial support where needed</Description><Identifier>_48f045aa-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>2.1.3</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Federation</Name><Description>Support the federation of user-supplied computing resources, testbeds, and sensors at the edge</Description><Identifier>_48f046f4-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>2.2</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The NAIRR should also support the federation of user-supplied computing resources, testbeds, and sensors at the edge.</OtherInformation></Objective></Goal><Goal><Name>Capacity &amp; Capability</Name><Description>Address both the capacity and capability needs of the AI research community</Description><Identifier>_48f04816-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>AI Research Community</Name><Description/></Stakeholder><OtherInformation>Capacity and Capability ~ When fully implemented, the NAIRR should address both the capacity (i.e., ability to support many users) and capability (i.e., ability to train the most resource-intensive AI models) needs of the AI research community.</OtherInformation><Objective><Name>Computational Resources</Name><Description>Provide a mix of computational resources with a range of central processing unit (CPU) and graphics processing unit (GPU) options</Description><Identifier>_48f0492e-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>3.1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>To meet existing capacity needs, the NAIRR should provide a mix of computational resources (i.e., on-premise and commercial cloud, dedicated, and shared resources) with a range of central processing unit (CPU) and graphics processing unit (GPU) options with multiple accelerators per node, high-speed networking, and sufficient memory capacity (i.e., at least one terabyte per node). The exact balance of computational resources will depend on the results of resource provider funding opportunities. Users should have the option of selecting which resources they would like to use through a range of mechanisms, including the user portal, direct command-line access, or optionally interactive “notebook”-like environments.</OtherInformation></Objective><Objective><Name>Supercomputer</Name><Description>Support at least one large-scale machine-learning supercomputer</Description><Identifier>_48f04a50-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>3.2</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>To meet users’ capability needs, the NAIRR system should include at least one large-scale machine-learning supercomputer capable of training 1 trillion-parameter models. This could be made available by leveraging an existing supercomputer or newly procured through a competitive bid process managed by the Operating Entity in consultation with the Steering Committee and relevant advisory boards.</OtherInformation></Objective></Goal><Goal><Name>Software Resources</Name><Description>Assess OSS packages most used by AI researchers</Description><Identifier>_48f04b68-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>AI Researchers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>TensorFlow</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>PyTorch</Name><Description/></Stakeholder><OtherInformation>NAIRR Software Resources ~ AI research has grown explosively through the development and dissemination of open source software (OSS) frameworks including TensorFlow, PyTorch, and their derivatives. Both these packages were developed by commercial entities and could have been kept proprietary. Instead, they were released as OSS projects, to the benefit of, and for further development by, the AI research community. The success of these projects has inspired many other OSS projects and tools.
^^
The Operating Entity, with advice from the Technology Advisory Board, should assess OSS packages most used by AI researchers and specify a standard software environment for the NAIRR federation.</OtherInformation><Objective><Name>Standard</Name><Description>Specify a standard software environment for the NAIRR federation</Description><Identifier>_48f04c8a-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>4.1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>This software environment should be containerized as a lightweight virtual machine, and be supported across resource providers. Academic teams with their own on-premise servers would be encouraged to adopt the NAIRR federation standard.</OtherInformation></Objective><Objective><Name>Workflows</Name><Description>Explore new AI workflow orchestration tools and templates for standard AI analysis tasks</Description><Identifier>_48f04db6-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>4.2</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>In addition, the Operating Entity should explore new AI workflow orchestration tools and templates for standard AI analysis tasks, such as cnvrg.io, which can meet the needs of industry researchers and might be suitable for adoption by the NAIRR federation.</OtherInformation></Objective></Goal><Goal><Name>Data &amp; Datasets</Name><Description>Provide a search and discovery service with metadata about the usage of all datasets</Description><Identifier>_48f04ed8-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>5</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Data and Datasets ~ The Operating Entity should provide a search and discovery service with metadata about the usage of all datasets. Such a service should be consistent with Section 202(c) of the Evidence Act.  It should be designed to dovetail with the capabilities anticipated through development of a Federal data catalog, but extend beyond Federal data.</OtherInformation><Objective><Name>Funding</Name><Description>Fund the creation of or provide continuing support to existing AI data repositories</Description><Identifier>_48f05018-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>5.1</SequenceIndicator><Stakeholder><Name/><Description/></Stakeholder><OtherInformation>The Operating Entity should support data resource providers by either funding the creation of or providing continuing support to existing AI data repositories.</OtherInformation></Objective><Objective><Name>Interoperability &amp; Competition</Name><Description>Publish interoperability guidelines for data repositories and encourage competion to become NAIRR data resource providers</Description><Identifier>_48f05158-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>5.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>NAIRR Data Resource Providers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Subcommittee on Open Science</Name><Description>National Science and Technology Council</Description></Stakeholder><OtherInformation>In coordination with the Technology Advisory Board, the Operating Entity should publish interoperability guidelines for such data repositories, and encourage data repositories to compete to become NAIRR data resource providers. These guidelines should be informed by the Desired Characteristics of Data Repositories for Federally Funded Research developed by the National Science and Technology Council’s Subcommittee on Open Science. Having such repositories and datasets visible, searchable, and discoverable inside the NAIRR, as well as implementing mechanisms to track dataset use, are important to the success of the NAIRR.</OtherInformation></Objective><Objective><Name>Security &amp; Federation</Name><Description>Federate computational, network, and data resources operating in accordance with security and access-control policies</Description><Identifier>_48f052b6-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>5.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>NAIRR-Open and NAIRR-Secure zones should federate computational, network, and data resources operating in accordance with security and access-control policies that are uniform within the zone, but different between zones, reflecting the restrictions associated with the data in each zone. NAIRR-Secure should coordinate and collaborate with the program office designated by the Office of Management and Budget to oversee the SAP, and others as appropriate, in making available and specifying security and user access controls required for restricted (confidential) government and third-party data. SAP is required by the Evidence Act to be the “front door” for accessing restricted data within the possession of Federal statistical agencies.</OtherInformation></Objective></Goal><Goal><Name>Datasets &amp; Metadata</Name><Description>Evaluate and characterize datasets into tiers, each with a different level of acceptance criteria</Description><Identifier>_48f053f6-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>6</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Dataset Acceptance Criteria and Metadata Standards ~ The Operating Entity should evaluate and characterize datasets into tiers, each with a different level of acceptance criteria. Examples include high, medium, and low levels of metadata; provenance; information about dataset usage, and the availability of persistent identifiers.</OtherInformation><Objective><Name>Evaluation &amp; Categorization</Name><Description>Ensure that each dataset is evaluated according to industry standards or best practices</Description><Identifier>_48f0554a-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>6.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>The Operating Entity should ensure that each dataset is evaluated according to industry standards or best practices and that a determination is made on how each should be categorized.</OtherInformation></Objective><Objective><Name>Data Catalog</Name><Description>Develop a Federal data catalog</Description><Identifier>_48f0568a-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>6.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Where possible, such cataloging efforts should be aligned with efforts to develop a Federal data catalog.</OtherInformation></Objective><Objective><Name>Formats</Name><Description>Provide a list of acceptable formats and leverage community-driven principles and standards</Description><Identifier>_48f057c0-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>6.3</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Research Data Alliance</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>NIST</Name><Description/></Stakeholder><OtherInformation>The Operating Entity should not define dataset standards, as this area continues to evolve rapidly and would be best addressed by the community of users. However, the Operating Entity should provide a public-facing list of acceptable formats to ensure compatibility with resources and tools, encourage broader use, and leverage existing community-driven principles and standards such as those developed by the Research Data Alliance and NIST, among others.</OtherInformation></Objective><Objective><Name>Documentation</Name><Description>Provide documentation with each directory or file</Description><Identifier>_48f05914-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>6.4</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Regardless of category, substantive documentation should be provided with each directory or file containing data.</OtherInformation></Objective><Objective><Name>Analysis Readiness</Name><Description>Specify the meaning of “analysis-ready” and categorize datasets accordingly</Description><Identifier>_48f05a54-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>6.5</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>The Operating Entity should also specify what it means for a dataset to be “analysis-ready” and categorize datasets accordingly. For example, an analysis-ready dataset should be in a structured format (e.g., a relational table or JSON34 or Neo4j35 formats) and should include details such as the semantics and provenance, information about the data-generation process, a data dictionary, related code, summary statistics for quality-assurance purposes, and information about how it has been used in previous analyses. Further, such a dataset should conform to standards in cases where datatypes are normally represented in a standard ontology (e.g., geographic information system [GIS] vector objects, gene ontology codes for molecules).</OtherInformation></Objective><Objective><Name>Rarity &amp; Importance</Name><Description>Transform important and rare datasets to become analysis-ready</Description><Identifier>_48f05b9e-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>6.6</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Not all datasets need be in analysis-ready form. Some types of data or partial datasets are important or rare, and can be contributed with the goal that others can help transform them into analysis-ready data.</OtherInformation></Objective></Goal><Goal><Name>Incentives &amp; Curation</Name><Description>Incentivize and curate contributed datasets and other resources</Description><Identifier>_48f05cf2-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Role of the Operating Entity in Incentivizing and Curating Contributed Datasets
and Other Resources</OtherInformation><Objective><Name>Data Service</Name><Description>Establish a data service that facilitates access to and additional use of existing curated datasets</Description><Identifier>_48f05e3c-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Since the quality of many AI models depends on high-quality training and test data, the Operating Entity should establish a data service that facilitates access to and additional use of existing curated datasets of value and interest to the NAIRR user community. Curation of AI data, models, tools, and workflows should be done by the user community in an AI data commons, facilitated by the NAIRR search and discovery platform. Such a community system, governed by terms of use as well as a review system, would facilitate data sharing and curation by members of the community.</OtherInformation></Objective><Objective><Name>Recognition &amp; Access</Name><Description>Incentivize researchers who contribute to the common good through high-profile NAIRR recognition and/or preferential access to NAIRR resources</Description><Identifier>_48f05f86-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7.2</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>In the context of a commons model, researchers who contribute to the common good through data curation and code sharing, and whose contributions are recognized and valued by relevant communities, could be incentivized through high-profile NAIRR recognition and/or preferential access to NAIRR resources.</OtherInformation></Objective><Objective><Name>Query, Discovery &amp; Curation</Name><Description>Test a service for searching for, discovering, and curating valuable external data as well as data generated with NAIRR resources</Description><Identifier>_48f060ee-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7.3</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>The NAIRR Operating Entity should test, on a trial basis, a service for searching for, discovering, and curating valuable external data as well as data generated with NAIRR resources.</OtherInformation></Objective><Objective><Name>AI Marketplaces</Name><Description>Contract with one or more commercial AI marketplaces to meet user data curation needs</Description><Identifier>_48f062f6-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7.3.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>One option would be to contract with one or more commercial AI marketplaces to meet its users' data curation needs. The “AI marketplace” is a powerful concept that has emerged in the commercial sector; it refers to the social and technical infrastructure through which the user community contributes, documents, and shares data, codes, and models. Contributions are validated and valued by the community, and community standards are enforced by the company managing the marketplace.</OtherInformation></Objective><Objective><Name>Data Commons</Name><Description>Develop an “AI data commons”</Description><Identifier>_48f06454-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7.3.2</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Another option is for the NAIRR to develop its own “AI data commons” with attributes similar to a commercial marketplace. Such an option is likely to be preferable for the federally funded NAIRR. However, since both commons and marketplace options have merit, the Operating Entity should have flexibility regarding development of data curation services, and the services should be implemented on a trial basis and evaluated for efficacy by the Operating Entity in the first five years of NAIRR operation.</OtherInformation></Objective><Objective><Name>Technical Support</Name><Description>Dedicate resources to technical support staff who can support community-driven curation efforts</Description><Identifier>_48f06620-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7.4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Users</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Contributors</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Curators</Name><Description/></Stakeholder><OtherInformation>Substantial Operating Entity resources should be dedicated to technical support staff who can support community-driven curation efforts. Data users, contributors, and curators will require support to understand and meet the technical standards of NAIRR data repositories.</OtherInformation></Objective><Objective><Name>Training &amp; Support</Name><Description>Provide training and additional support to ensure the integrity and quality of NAIRR datasets as well as to protect privacy, civil rights, and civil liberties</Description><Identifier>_48f0677e-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>7.5</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Further, training and additional support will be critical to the integrity and quality of NAIRR datasets, and to protect privacy, civil rights, and civil liberties.</OtherInformation></Objective></Goal><Goal><Name>Federal Data</Name><Description>Make federal datasets accessible</Description><Identifier>_48f068dc-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>The NAIRR and Existing Federal Government Data ~ Federal agencies hold data that could fuel foundational, use-inspired, and translational AI research in domains such as transportation, healthcare, and natural hazards research. Sources of Federal agency data include statistical data, administrative data, and data from federally funded intramural and extramural research. While some of these datasets are already accessible to the public, many others are not.
^^
Since Federal datasets could be highly valuable to AI research and advance national goals, there are three other Federal Government data efforts with which the NAIRR could engage.</OtherInformation><Objective><Name>data.gov</Name><Description>Work with data.gov to encourage additional contributions</Description><Identifier>_48f06a62-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>data.gov</Name><Description/></Stakeholder><OtherInformation>One is data.gov, which is a website that points to other resources containing information and data generated by agency or agency-funded projects. Most of the retrievable data on data.gov are in web or text form, which might be of interest to some NAIRR researchers. However, scientific numerical datasets are deeply buried in data.gov and not easily accessible. The Operating Entity and Program Management Office could work with data.gov to encourage additional contributions conforming to NAIRR data acceptance criteria, which should include measures of data use.</OtherInformation></Objective><Objective><Name>SAP</Name><Description>Enable discovery and access to restricted data acquired by Federal statistical agencies through a single application process and portal</Description><Identifier>_48f06bca-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8.2</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>SAP</Name><Description/></Stakeholder><OtherInformation>Another is the SAP, through which researchers will be able to discover and apply for access to restricted data acquired by Federal statistical agencies through a single application process and portal.</OtherInformation></Objective><Objective><Name>Acquisition, Linkage &amp; Protection</Name><Description>Provide additional capability for data acquisition, linkage, and protection</Description><Identifier>_48f06d3c-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Finally, the National Secure Data Service (NSDS) demonstration project, established by the CHIPS and Science Act of 2022, has the potential to complement the SAP and existing statistical agency efforts with additional capability for data acquisition, linkage, and protection (see Box 6 for more details). </OtherInformation></Objective><Objective><Name>Statistical Policy</Name><Description>Establish a NAIRR-Federal Interagency Council on Statistical Policy (ICSP) working group</Description><Identifier>_48f06ec2-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8.4</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Interagency Council on Statistical Policy (ICSP)</Name><Description/></Stakeholder><OtherInformation>The Steering Committee should facilitate the establishment of a NAIRR-Federal Interagency Council on Statistical Policy (ICSP) working group.</OtherInformation></Objective><Objective><Name>Secure Node</Name><Description>Assess options for establishing a secure node for the purpose of enabling large-scale AI analysis of government data for statistical purposes</Description><Identifier>_48f07034-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8.4.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>This working group should collaborate to assess options for establishing a secure node for the purpose of enabling large-scale AI analysis of government data for statistical purposes.</OtherInformation></Objective><Objective><Name>Protocols &amp; Controls</Name><Description>Define the Confidential Information Protection and Statistical Efficiency Act-compliant data access protocols and controls</Description><Identifier>_48f071a6-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8.4.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Where such resources are not intended to be made accessible via the SAP or the NSDS demonstration project, the working group should define the Confidential Information Protection and Statistical Efficiency Act-compliant data access protocols and controls.  This NAIRR-ICSP collaboration should facilitate the provisioning of timely access for appropriate (i.e., approved) projects to restricted (i.e., confidential) government and third-party data.</OtherInformation></Objective><Objective><Name>State &amp; Local Datasets</Name><Description>Encourage and support additional contributions of State and local datasets</Description><Identifier>_48f07336-cf63-11ed-b3e1-ed4b1883ea00</Identifier><SequenceIndicator>8.5</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>State Government Agencies</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Local Government Agencies</Name><Description/></Stakeholder><OtherInformation>The NAIRR should also encourage and support additional contributions of State and local datasets conforming to NAIRR data acceptance criteria, and subject to the legal requirements of the State and local government agencies, either by working with data.gov or the eventual NSDS.


In terms of existing high-quality data repositories managed by agencies such as the National Institutes of Health (NIH) and NASA, the Operating Entity will need to determine whether to reproduce large datasets that are already available from these other sources or find other means of coordinating access for NAIRR researchers. This coordination could benefit from regular convening of leadership from various Federal data efforts to identify ways to improve coordination and avoid inefficiency or redundancy.</OtherInformation></Objective></Goal></StrategicPlanCore><AdministrativeInformation><StartDate>2023-01-31</StartDate><EndDate/><PublicationDate>2023-04-04</PublicationDate><Source>https://www.ai.gov/wp-content/uploads/2023/01/NAIRR-TF-Final-Report-2023.pdf</Source><Submitter><GivenName>Owen</GivenName><Surname>Ambur</Surname><PhoneNumber/><EmailAddress>Owen.Ambur@verizon.net</EmailAddress></Submitter></AdministrativeInformation></PerformancePlanOrReport>