<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<?xml-stylesheet type="text/xsl" href="../part2stratml.xsl"?><PerformancePlanOrReport><Name>The Chief Data Officer in Government: A CDO Playbook</Name><Description>The CDO Playbook, produced by Georgetown University’s Beeck Center and Deloitte’s Center for Government Insights, explores some of the hardest questions facing CDOs today. The playbook draws on conversations we’ve had over the past year with CDOs from multiple levels of government as well as in the private, nonprofit, and social sectors. Insights from these leaders shed light on opportunities and potential growth areas for the use of data and the role of CDOs within government.</Description><OtherInformation>We hope this playbook will help catalyze the further evolution of CDOs within government and provide an accessible guide for executives who are still evaluating the creation of these positions.</OtherInformation><StrategicPlanCore><Organization><Name>Deloitte Center for Government Insights</Name><Acronym>DC4GI</Acronym><Identifier>_1f49c4a6-3105-11ea-a45d-83222a83ea00</Identifier><Description>The Deloitte Center for Government Insights shares inspiring stories of government innovation, looking at what’s behind the adoption of new technologies and management practices. We produce cutting-edge research that guides public officials without burying them in jargon and minutiae, crystalizing essential insights in an easy-to-absorb format. Through research, forums, and immersive workshops, our goal is to provide public officials, policy professionals, and members of the media with fresh insights that advance an understanding of what is possible in government transformation.</Description><Stakeholder StakeholderTypeType="Organization"><Name>Beeck Center for Social Impact + Innovation</Name><Description>The Beeck Center for Social Impact + Innovation at Georgetown University engages global leaders to drive social change at scale. Through our research, education, and convenings, we provide innovative tools that leverage the power of capital, data, technology, and policy to improve lives. We embrace a cross-disciplinary approach to building solutions at scale.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Sonal Shah</Name><Description>Co-Author -- Sonal Shah is executive director, professor of practice, at the Beeck Center for Social Impact + Innovation. She is based in Washington, DC.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>William D. Eggers</Name><Description>Co-Author -- William D. Eggers is the executive director of Deloitte’s Center for Government Insights, where he is responsible for the firm’s public sector thought leadership. He is based in Arlington, VA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Executives</Name><Description>The playbook is written for government executives as well as for government CDOs. For executives, it provides an overview of the types of functions that CDOs across the country are performing.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Chief Data Officers</Name><Description>For CDOs, it offers a guide to understanding the trends affecting public sector data, and provides practical guidance on strategies they can pursue for effective results.</Description></Stakeholder></Organization><Vision><Description/><Identifier>_1f49c640-3105-11ea-a45d-83222a83ea00</Identifier></Vision><Mission><Description>To catalyze the evolution of CDOs within government and provide a guide for executives who are evaluating the creation of these positions.</Description><Identifier>_1f49c758-3105-11ea-a45d-83222a83ea00</Identifier></Mission><Value><Name>Prioritization</Name><Description>Why is data important?</Description></Value><Value><Name>Transparency</Name><Description>Public demand for transparency and accountability</Description></Value><Value><Name>Accountability</Name><Description/></Value><Value><Name>Access</Name><Description>Increased access to large amounts of data</Description></Value><Value><Name>Security</Name><Description>Responsibility for data security</Description></Value><Value><Name>Innovation</Name><Description>Technology innovation and exponential disruptors driving added complexity</Description></Value><Value><Name>Needs</Name><Description>Changing citizen needs and preferences</Description></Value><Value><Name>Preferences</Name><Description/></Value><Value><Name>Efficiency</Name><Description>Budget constraints driving the need for greater operational efficiency</Description></Value><Value><Name>Responsibility</Name><Description>Responsibility to limit fraud, waste, and abuse</Description></Value><Value><Name>Helpfulness</Name><Description>Where can data help?* Effectiveness: “Do what we do better”* Efficiency: “Do more with less”* Fraud, waste, and abuse: “Find and prevent leakage”* Transparency and citizen engagement: “Build trust”</Description></Value><Value><Name>FAIR Principles</Name><Description>WHAT IS FAIR?3The FAIR principles are a set of guidingprinciples for scientific data managementand stewardship to support innovationand discovery. Distinct from peer initiativesthat focus on the human scholar, theFAIR principles put specific emphasison enhancing the ability of machinesto automatically find and use data—inother words, making data “machine-actionable”—in addition to supporting itsreuse by individuals. Widely recognized andsupported in the scientific community, theprinciples posit that data should be:• Findable. Data must have uniqueidentifiers that effectively label it withinsearchable resources.• Accessible. Data must be easilyretrievable via open systems that haveeffective and secure authentication andauthorization procedures.• Interoperable. Data should “use andspeak the same language” by usingstandardized vocabularies.• Reusable. Data must be adequatelydescribed to a new user, include clearinformation about data usage licenses,and have a traceable “owner’s manual”or provenance.</Description></Value><Goal><Name>Storytelling</Name><Description>Connect data to residents through data storytelling</Description><Identifier>_1f49c8ca-3105-11ea-a45d-83222a83ea00</Identifier><SequenceIndicator>1</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>William D. Eggers</Name><Description>Co-Author -- William D. Eggers is the executive director of Deloitte’s Center for Government Insights, where he is responsible for the firm’s public sector thought leadership.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Amrita Datar</Name><Description>Co-Author -- Amrita Datar is an assistant manager with the Deloitte Center for Government Insights.</Description></Stakeholder><OtherInformation>The story is mightier than the spreadsheet -- In the past decade, many governments have taken significant strides in the open data movement by making thousands of data sets available to the general public. But simply publishing a data set on an open data portal or website is not enough. For data to have the most impact, it’s essential to turn those lines and dots on a chart or numbers in a table into something that everyone can understand and act on.Data itself is often disconnected from the shared experiences of the American people. An agency might collect and publish data on a variety of areas, but without the context of how it impacts citizens, it might not be as valuable. So how do we connect data to the citizenry’s shared everyday lives? Through a language that is deeply tied to our human nature—stories...Four ways to harness the power of data stories [are documented as objectives below].LOOKING AHEAD -- The value of data is determined not by the data itself, but by the story it tells and the actions it empowers us to take. But for citizens to truly feel connected to data, they need to see more than just numbers on a page; they need to understand what those numbers really mean for them. To make this connection, CDOs and data teams will need to invest in how they present data and think creatively about new formats and platforms.</OtherInformation><Objective><Name>Visualization</Name><Description>SHOW, DON’T JUST TELL</Description><Identifier>_1f49c9ec-3105-11ea-a45d-83222a83ea00</Identifier><SequenceIndicator>1.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Washington, DC</Name><Description>In Washington, DC, the interactive websiteDistrict Mobility turns data on the DC area’s multimodaltransportation system into map-based visualstories. Which bus routes serve the most riders?How do auto travel speeds vary by day of week andtime of day on different routes? How punctual is the bus service in differentareas of town? These are justsome of the questions thatresidents and city plannerscan find answers to throughthe site.2</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>DataUSA</Name><Description>Similarly, DataUSAcombines publicly accessibleUS government datafrom a variety of agenciesand brings it to life in morethan 2 million visualizations.In addition to allowing usersto search, map, compare, anddownload data sets, the site alsoshows them what kinds of insightsthe data can reveal through “DataUSA stories.”“People do not understand the world by looking atnumbers; they understand it by looking at stories,”says Cesar Hidalgo, director of the MassachusettsInstitute of Technology Media Lab’s Macro Connectionsgroup, and one of DataUSA’s creators.3DataUSA’s stories combine maps, charts, and othervisualizations with narratives around a range oftopics—from the damage done by opioid addictionto real estate in the rust belt to income inequality inAmerica—that might pique citizens’ interest. Somestates, such as Virginia, have also embedded interactivecharts from DataUSA into their economicdevelopment portals.</Description></Stakeholder><OtherInformation>As human beings, our brains are wired toprocess visual data better than other forms of data.In fact, the human brain processes images 60,000times faster than text.1 For example, public healthdata shown on a map might be infinitely moremeaningful and accessible to citizens than a heavytable with the same information. Increasingly,governments are tapping into the power of datavisualization to connect with citizens.</OtherInformation></Objective><Objective><Name>Impact</Name><Description>PICK A HIGH-IMPACT PROBLEM</Description><Identifier>_1f49cafa-3105-11ea-a45d-83222a83ea00</Identifier><SequenceIndicator>1.2</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>New Orleans</Name><Description>In the aftermath of hurricane Katrina, manyneighborhoods across New Orleans were full ofblighted and abandoned buildings—more than40,000 of them. Residents and city staff couldn’teasily get information on the status of blightedproperties—data that was necessary for communitiesto come together and make decisions aroundrebuilding their neighborhoods.4</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Code for America</Name><Description>New Orleans city staff worked with a team ofCode for America fellows to build an open data-poweredweb application called Blight Status, whichenabled anyone to look up an address and see whatreports had been made on the property—blightreports, inspections, hearings, and scheduled demolitions.The app connected both citizens and citybuilding inspectors to the data and presented it inan easily accessible map-based format along withthe context needed to make it actionable.5</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Black Women</Name><Description>Data-driven stories can also reveal hidden truthsabout institutionalized biases. Across America,social justice movements are highlighting citizendisparities. While protests grab the attention ofsome and repel others, telling compelling storiessupported by data can spur meaningful shifts inthinking and outcomes. For example, in the UnitedStates, the data shows that black women are 243percent more likely to die than white women frombirth-related complications.6 This disparity persistsfor black women who outpace white women in educationlevel, income, and access to health care. Thedata challenges an industry to address the quality ofcare provided to this population of Americans.</Description></Stakeholder><OtherInformation>To connect with citizens across groups, focusdata and storytelling efforts around issues that havea far-reaching impact on their lives.</OtherInformation></Objective><Objective><Name>Decision-Making</Name><Description>SHARE HOW DATA DRIVES DECISION-MAKING</Description><Identifier>_1f49cc26-3105-11ea-a45d-83222a83ea00</Identifier><SequenceIndicator>1.3</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Kansas City</Name><Description>Consider the example of Kansas City’s KCStatprogram. KCStat meetings are held each monthto track the city’s progress toward its goals. Datais used to drive the conversation around a host ofissues, from public safety and community health toeconomic development and housing. Citizens areinvited to the meetings, and stats and highlightsfrom meetings are even shared on Twitter (#KCStat)to encourage participation and build awareness.7“As data becomes ingrained systemically in youroperation, you can use facts and data to create,tweak, sustain, and perfect programs that willprovide a real benefit to people, and it’s verifiable bythe numbers,” said Kansas City mayor Sly James inan interview for Bloomberg’s What Works Cities.8The city also publishes a blog called Chartlandthat tells stories drawn from the city’s data. Somefocus on themes from KCStat meetings, whileothers, often written by the city’s chief data officer(CDO) and the office of the city manager, explorepertinent city issues such as the risk of lead poisoningin older homes, patterns in 311 data, or howresults from a citizen satisfaction survey helpeddrive an infrastructure repair plan.9 These blogs areconversational and easy to understand, helping tohumanize data that can seem intimidating to many.</Description></Stakeholder><OtherInformation>Another way to bring citizens closer to data thatmatters to them is by telling the story of how thatdata can shape government decisions that impacttheir lives. This can be accomplished through a blog,a talk, a case study, or simply in the way public officialscommunicate successes to their constituents.</OtherInformation></Objective><Objective><Name>Interaction</Name><Description>MAKE STORYTELLING A TWO-WAY STREET</Description><Identifier>_1f49cd48-3105-11ea-a45d-83222a83ea00</Identifier><SequenceIndicator>1.4</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>New York City</Name><Description>To celebrate the five-year anniversary of the NYCOpen Data Law, for instance, New York City’s OpenData team organized its first-ever Open Data Weekin 2017. The week’s activities included 12 eventsrevolving around open data, which attracted over900 participants. The city’s director of open dataalso convened “Open Data 101,” a training sessiondesigned specifically to teach nontechnical usershow to work with open data sets.10</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Pittsburgh</Name><Description>It’s important for event organizers to be cognizantthat, for data storytelling to bring citizenscloser to data, activities should be designed toenable participation for all—not just those whoare already skilled with technology and data. Forexample, when Pittsburgh hosted its own Open DataDay—an all-day drop-in event for citizens to engagein activities around data—the event included a lowtech“Dear Data” project in which participants couldhand-draw a postcard to tell a data-based story.Organizers also stipulated that activity facilitatorsshould adopt a “show and play” format—a demofollowed by a hands-on activity instead of a staticpresentation—to encourage open conversation andparticipation.11</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Ben Wellington</Name><Description>Citizens telling their own stories with data canshed light on previously unknown challenges andopportunities, giving them a voice to drive change.For example, Ben Wellington, a data enthusiastlooking through parking violation data in NewYork City, discovered millions of dollars’ worthof erroneous tickets issued for legally parkedcars.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Patrol Officers</Name><Description>Some patrol officers were unfamiliar with arecent parking law change and continued to issuetickets—a problem that the city has since corrected,thanks to Wellington’s analysis.12</Description></Stakeholder><OtherInformation>Hackathons and open data-themed events givecitizens a way to engage with data sets in guided settingsand learn to tell their own stories with the data.</OtherInformation></Objective></Goal><Goal><Name>Obstacles</Name><Description>Overcome obstacles to open data-sharing</Description><Identifier>_1f49ce7e-3105-11ea-a45d-83222a83ea00</Identifier><SequenceIndicator>2</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Adam Neufeld</Name><Description>Author -- Adam Neufeld is a senior fellow at the Beeck Center for Social Impact + Innovation. He is based in Washington, DC.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>CDOs looking to unleash the potential of opendata should consider ways that they could addressthese obstacles.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>General Services Administration</Name><Description>One potential approach is to centralizedecision-making authority and technicalcapabilities rather than having these distributedamong the numerous offices and departments that“own” the data. The General Services Administration,for example, created a chief data officer position toact in this capacity.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Congress</Name><Description>Several other agencies havedone the same, and Congress is currently consideringlegislation to require every agency to do so.10</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Open Data Community</Name><Description>The open data community, for its part, can playan important part in encouraging data-sharing byhelping agencies understand what data would be most useful under what conditions. CDOs sometimesdo not have the political strength or themanagement or technical bandwidth to release allof their agencies’ data, even if this were always desirable,so prioritization is key.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Regulated Entities</Name><Description>Regulated entitiesand beneficiaries should also help the governmentdetermine what the next-best alternative is if fullopenness is not possible.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Beneficiaries of Regulations</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>US Department of Health and Human Services</Name><Description>A few agencies, such asthe US Department of Health and Human Serviceswith its Demand-Driven Open Data effort, haveinvited the public to engage in prioritization. Topromote greater openness, however, such effortsshould be spread across more agencies and involvemore levels at those agencies.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Officials</Name><Description>Understanding theperspectives of those outside government can helpofficials balance the trade-off between releasingdata and controlling the risks and costs.CDOs’ leadership will be important in encouraginggovernment to move swiftly to release allappropriate data that could benefit our society, democracy,and economy.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Private Sector</Name><Description>To be most effective, theymay need private-sector input and policy guidancethat can help them and support them on the opendata journey.</Description></Stakeholder><OtherInformation>How CDOs can overcome obstacles to open data-sharing -- OPEN DATA HAS been a hot topic in government for the past decade. Various politicians from across the spectrum have extolled the benefits of increasing access to and use of government data, citing everything from enhanced transparency to greater operating efficiency.1While the open data movement seems to have achieved some successes, including the DATA Act2 and data.gov,3 we have yet to achieve the full potential of open data. The McKinsey Global Institute, for example, estimates that opening up more data could result in more than $3 trillion in economic benefits.4It is time for the open data community to pivot based on the lessons learned over the past decade, and governmental chief data officers (CDOs) can lead the way.  Much valuable government data remains inaccessible to the public. In some cases, this is because the data includes personally identifiable information.But in other situations, data remains unshared because government has procured a proprietary system that prevents sharing. Moreover, when government does share data, it sometimes does so in spreadsheets or in other formats that can limit its usefulness, rather than in a format such as an application programming interface (API) that would allow for easier use. In fact, some of the potentially most valuable public information, such as financial regulatory filings, is typically not machine-readable.  CDOs looking to achieve greater benefits through open data should devise a plan that addresses both the technical and administrative challenges of data sharing, including:</OtherInformation><Objective><Name>Incentives</Name><Description>Align incentives between political leaders and their staff.</Description><Identifier>_1f49cfc8-3105-11ea-a45d-83222a83ea00</Identifier><SequenceIndicator>2.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Principals</Name><Description>The principals (here, the public, legislators, and, to some extent, executive branch leaders) generally want data to be open because they stand to reap the societal and/or reputational benefits of whatever comes from releasing it.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Agents</Name><Description>However, the decision of whether to standardize or release data is made by an agent (here, usually some combination of program managers, information technology professionals, and lawyers). The agent tends to gain little direct benefit from releasing the data — but they could face substantial costs in doing so.Not only would they need to do the hard work of standardization, but they would incur the risk of reputational damage, stress, or termination if the data they release turns out to be inaccurate, creates embarrassment for the program, or compromises privacy, national security, or business interests. As a result, even if a political leader wants to share data, there may still be obstacles to doing so.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Program Managers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Information Technology Professionals</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Lawyers</Name><Description/></Stakeholder><OtherInformation>Mismatched incentives between political leaders and their staff: Not all data can or should be shared publicly. Agencies are prohibited from sharing personally identifiable data, medical data, and certain other information. There are, however, many gray areas regarding what can or cannot be disclosed. In these instances, the decision on whether and how to standardize or publish a government data set has all the ingredients of a standard principal-agent problem in economics.5</OtherInformation></Objective><Objective><Name>Sharing</Name><Description>Consider intermediate data-sharing options that could provide much of the benefit of full disclosure, but at less cost and/or lower risk.</Description><Identifier>_6ffd1a7e-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>2.2</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Center for Medicare and Medicaid Services</Name><Description>For example, the Center forMedicare and Medicaid Services allows companiesto apply for limited, secure access totransaction data to help them develop productsthat aim to improve health outcomes or reducehealth spending.7</Description></Stakeholder><OtherInformation>An “all or nothing” approach to data-sharing:The discussion of open data is oftenpresented in binary terms: Either data is open,meaning that it is publicly available in a standardizedformat for download on a website, orit is not accessible to outsiders at all. This typeof thinking takes intermediate options off thetable that could provide much of the benefit offull disclosure, but at less cost and/or lower risk.The experience of federal statistical agenciessuggests that intermediate approaches couldallow even some sensitive data to be shared ona limited basis.6</OtherInformation></Objective><Objective><Name>Expertise</Name><Description>Acquire required technical expertise.</Description><Identifier>_6ffd1d6c-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>2.3</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Lack of technical expertise: Releasing adata set is generally time-consuming technicalwork that may require cleaning the data anddeciding on privacy protections. Some governmentsmay have limited in-house technologicalexpertise, however, and these technical expertsare often needed for other competing priorities.The skills needed to appropriately releasedata sets that contain sensitive information areeven more technical, requiring people with anunderstanding of advanced cryptographic andtechnical approaches such as synthetic data8and secure multiparty computation.9 Usually,the subject-matter experts who control whethera given data set will be opened do not have thisexpertise. This is understandable, as such skillswere not historically necessary or even useful,but the skill set gap can prevent governmentsfrom sharing data even when all stakeholdersagree that it should be shared.</OtherInformation></Objective><Objective><Name>Prioritization</Name><Description>Prioritize data sets.</Description><Identifier>_6ffd1f24-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>2.4</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Difficulty in prioritizing data sets: Just asreleasing data typically requires a rare combinationof subject matter and technical expertise,so can figuring out which data sets to prioritize.How government data might be put to beneficialuse requires imagination from peoplewith varied perspectives. Government officialscannot always predict what data sets, especiallywhen used in concert with other data sets, mightprove transformative. This is even more truewhen considering the details of how data shouldbe shared.</OtherInformation></Objective></Goal><Goal><Name>Machine Learning</Name><Description>Promote machine learning in government.</Description><Identifier>_6ffd2334-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>3</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>David Schatsky</Name><Description>Co-Author -- David Schatsky, a director with Deloitte Services LP, analyzes emerging technology and business trends for Deloitte’s leaders and clients. He is based in New York City.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Rameeta Chauhan</Name><Description>Co-Author -- Rameeta Chauhan, of Deloitte Services India Pvt. Ltd., tracks and analyzes emerging technology and business trends, with a primary focus on cognitive technologies, for Deloitte’s leaders and clients. She is based in Mumbai, India.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>Prepare for the mainstreaming of machine learning -- Collectively, the five vectors of machine learningprogress can help reduce the challenges governmentagencies may face in investing in machine learning.They can also help agencies already using machinelearning to intensify their use of the technology. Theadvancements can enable new applications acrossgovernments and help overcome the constraints oflimited resources, including talent, infrastructure,and data to train the models.CDOs have the opportunity to automate some ofthe work of often oversubscribed data scientists andhelp them add even more value. A few key thingsagencies should consider are:• Ask vendors and consultants how they use datascience automation.• Keep track of emerging techniques such as datasynthesis and transfer learning to ease the challengeof acquiring training data.• Investigate whether the agency’s cloud providersoffer computing resources that are optimized formachine learning.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Companies</Name><Description>Onerecent survey of 3,100 executives from small,medium, and large companies across 17 countriesfound that fewer than 10 percent of companies wereinvesting in machine learning.2</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Machine Learning Practitioners</Name><Description>A number of factors are restraining the adoptionof machine learning in government and theprivate sector. Qualified practitioners are in shortsupply.3 Tools and frameworks for doing machinelearning work are still evolving.4 It can be difficult,time-consuming, and costly to obtain large datasetsthat some machine learning model-developmenttechniques require.5</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Executives</Name><Description>Then there is the black box problem. Even whenmachine learning models can generate valuableinformation, many government executives seem reluctantto deploy them in production. Why? In part,possibly because the inner workings of machinelearning models are inscrutable, and some peopleare uncomfortable with the idea of running their operationsor making policy decisions based on logicthey don’t understand and can’t clearly describe.6</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Officials</Name><Description>Other government officials may be constrained byan inability to prove that decisions do not discriminateagainst protected classes of people.7 Using AIgenerally requires understanding all requirementsof government, and it requires making the blackboxes more transparent.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Protected Classes of People</Name><Description/></Stakeholder><OtherInformation>How CDOs can promote machine learning in government -- ARTIFICIAL INTELLIGENCE (AI) holds tremendouspotential for governments, especiallymachine learning technology, which canhelp discover patterns and anomalies and makepredictions. There are five vectors of progress thatcan make it easier, faster, and cheaper to deploymachine learning and bring the technology intothe mainstream in the public sector. As the barrierscontinue to fall, chief data officers (CDOs) haveincreasing opportunities to begin exploring applicationsof this transformative technology.Current obstacles -- Machine learning is one of the most powerfuland versatile information technologies availabletoday.1 But most organizations, even in the privatesector, have not begun to use its potential...Progress in these five areascan help overcome barriers toadoption -- There are five vectors of progress in machinelearning that could help foster greater adoptionof machine learning in government (see figure 1).Three of these vectors include automation, datareduction, and training acceleration, which makemachine learning easier, cheaper, and/or faster.The other two are model interpretability and localmachine learning, both of which can open up applicationsin new areas.</OtherInformation><Objective><Name>Data Science</Name><Description>AUTOMATE DATA SCIENCE</Description><Identifier>_6ffd258c-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>3.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Scientists</Name><Description>Automating these tasks canmake data scientists in governmentmore productive andmore effective.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Airbnb</Name><Description>For instance,while building customerlifetime value models forguests and hosts, data scientistsat Airbnb used anautomation platform to test multiplealgorithms and design approaches, w h i c hthey would not likely have otherwise had the timeto do. This enabled Airbnb to discover changes itcould make to its algorithm that increased the algorithm’saccuracy by more than 5 percent, resultingin the ability to improve decision-making and interactionswith the Airbnb community at very granularlevels.9</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Companies</Name><Description>A growing number of tools and techniques fordata science automation, some offered by establishedcompanies and others by venture-backedstartups, can help reduce the time required toexecute a machine learning proof of concept frommonths to days.10 And automating data sciencecan mean augmenting data scientists’ productivity,especially given frequent talent shortages. As theexample above illustrates, agencies can use datascience automation technologies to expand theirmachine learning activities.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Startups</Name><Description/></Stakeholder><OtherInformation>Developing machine learning solutions requiresskills primarily from the discipline of data science,an often-misunderstood field. Data science can beconsidered a mix of art and science—and digitalgrunt work. Almost 80 percent of the work thatdata scientists spend their time on can be fully orpartially automated, giving them time to spend onhigher-value issues.8 This includes data wrangling—preprocessing and normalizing data, filling inmissing values, or determining whether to interpretthe data in a column as a number or a date; exploratorydata analysis—seeking to understand thebroad characteristics of the data to help formulatehypotheses about it; feature engineering and selection—selecting the variables in the data that are most likely correlated with whatthe model is supposed to predict;and algorithm selection andevaluation—testing potentiallythousands of algorithms to assesswhich ones produce the most accurateresults.</OtherInformation></Objective><Objective><Name>Training Data</Name><Description>REDUCE THE NEED FOR TRAINING DATA</Description><Identifier>_6ffd274e-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>3.2</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Deloitte LLP</Name><Description>A Deloitte LLP team tested a tool that made itpossible to build an accurate machine learningmodel with only 20 percent of the training datapreviously required by synthesizing the remaining80 percent. The model’s task was to analyze jobtitles and job descriptions—which are often highlyinconsistent in large organizations, especiallythose that have grown by acquisition—and thencategorize them into a more consistent, standardset of job classifications. To learn how to do this,the model needed to be trained through exposureto a few thousand accurately classified examples.Instead of requiring analysts to laboriously classify(“label”) these thousands of examples by hand, thetool made it possible to take a set of labeled datajust 20 percent as large and automatically generatea fuller training dataset. And the resulting dataset,composed of 80 percent synthetic data, trainedthe model just as effectively as a hand-labeled realdataset would have.Synthetic data can not only make it easier toget training data, but also make it easier for organizationsto tap into outside data science talent.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>MIT</Name><Description>A number of organizations have successfullyengaged third parties or used crowdsourcing todevise machine learning models, posting their datasetsonline for outside data scientists to work with.12This can be difficult, however, if the datasets are proprietary. To address this challenge, researchersat MIT created a synthetic dataset that they thenshared with an extensive data science community.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Scientists</Name><Description>Data scientists within the community built machinelearning models using the synthetic data. In 11 outof 15 tests, the models developed from the syntheticdata performed as well as those trained on realdata.13</Description></Stakeholder><OtherInformation>Developing machine learning models typicallyrequires millions of data elements. This can be amajor barrier, as acquiring and labeling data can betime-consuming and costly. For example, a medicaldiagnosis project that requires MRI images labeledwith a diagnosis requires a lot of images and diagnosesto create predictive algorithms. It can costmore than $30,000 to hire a radiologist to reviewand label 1,000 images at six images an hour. Additionally,privacy and confidentiality concerns,particularly for protected data types, canmake working with data more time-consumingor difficult.A number of potentially promising techniquesfor reducing the amount of training datarequired for machine learning are emerging. Oneinvolves the use of synthetic data, generated algorithmicallyto create a synthetic alternative to mimicthe characteristics of real data.11 This technique hasshown promising results...Another technique that could reduce the needfor extensive training data is transfer learning. Withthis approach, a machine learning model is pretrainedon one dataset as a shortcut to learning anew dataset in a similar domain such as languagetranslation or image recognition. Some vendorsoffering machine learning tools claim their use oftransfer learning has the potential to cut the numberof training examples that customers need to provideby several orders of magnitude.14</OtherInformation></Objective><Objective><Name>Acceleration</Name><Description>EVOLVE TECHNOLOGY FOR ACCELERATED LEARNING</Description><Identifier>_6ffd2b86-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>3.3</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Microsoft</Name><Description>For instance, a Microsoftresearch team, using GPUs, completed a system thatcould recognize conversational speech as capably ashumans in just one year. Had the team used onlyCPUs, according to one of the researchers, the sametask would have taken five years.16</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Google</Name><Description>Google has statedthat its own AI chip, the Tensor Processing Unit(TPU), when incorporated into a computing systemthat also includes CPUs and GPUs, provided such aperformance boost that it helped the company avoidthe cost of building a dozen extra data centers.17 Thepossibility of reducing the cost and time involved inmachine learning training could have big implicationsfor government agencies, many of which havea limited number of data scientists.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Early Adopters</Name><Description>Early adopters of these specialized AI chipsinclude some major technology vendors and researchinstitutions in data science and machinelearning, but adoption also seems to be spreading tosectors such as retail, financial services, and telecom.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Retail Sector</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Financial Services Sector</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Telecom Sector</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Cloud Providers</Name><Description>With every major cloud provider—including IBM,Microsoft, Google, and Amazon Web Services—offeringGPU cloud computing, accelerated trainingwill likely soon become available to public sectordata science teams, making it possible for them tobe fast followers. This would increase these teams’productivity and allow them to multiply the numberof machine learning applications they undertake.18</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>IBM</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Microsoft</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Google</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Amazon Web Services</Name><Description/></Stakeholder><OtherInformation>Because of the large volumes of data and complexalgorithms involved, the computational processof training a machine learning model can take along time: hours, days, even weeks.15 Only thencan the model be tested and refined. Now, somesemiconductor and computer manufacturers—bothestablished companies and startups—are developingspecialized processors such as graphics processingunits (GPUs), field-programmable gate arrays, andapplication-specific integrated circuits to slash thetime required to train machine learning models byaccelerating the calculations and by speeding up thetransfer of data within the chip.These dedicated processors can help organizationssignificantly speed up machine learningtraining and execution, which in turn could bringdown the associated costs. </OtherInformation></Objective><Objective><Name>Transparency</Name><Description>MAKE RESULTS TRANSPARENT</Description><Identifier>_6ffd2dfc-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>3.4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Physicians</Name><Description>Physicians and business leaders, forinstance, may not accept a medical diagnosis orinvestment decision without a credible explanationfor the decision. In some cases, regulations mandatesuch explanations.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Business Leaders</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>MIT Researchers</Name><Description>Techniques are emerging that can help shinelight inside the black boxes of certain machinelearning models, making them more interpretableand accurate. MIT researchers, for instance,have demonstrated a method of training a neuralnetwork that delivers both accurate predictions andrationales for those predictions.19 Some of thesetechniques are already appearing in commercialdata science products.20</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>As it becomes possible to build interpretablemachine learning models, government agenciescould find attractive opportunities to use machinelearning. Some of the potential application areasinclude child welfare, fraud detection, and diseasediagnosis and treatment.21</Description></Stakeholder><OtherInformation>Machine learning models often suffer from theblack-box problem: It is impossible to explain withconfidence how they make their decisions. Thiscan make them unsuitable or unpalatable for manyapplications. </OtherInformation></Objective><Objective><Name>Deployment</Name><Description>DEPLOY LOCALLY</Description><Identifier>_6ffd30d6-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>3.5</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Microsoft Research Lab</Name><Description>Microsoft Research Lab’s compressionefforts resulted in models that were 10 to 100 timessmaller than earlier models.24 On the hardware end,various semiconductor vendors have developed orare developing their own power-efficient AI chips tobring machine learning to mobile devices.25</Description></Stakeholder><OtherInformation>The emergence of mobile devices as a machinelearning platform is expanding the number of potentialapplications of the technology and inducingorganizations to develop applications in areas suchas smart homes and cities, autonomous vehicles,wearable technology, and the industrial Internet ofThings.The adoption of machine learning will growalong with the ability to deploy the technology whereit can improve efficiency and outcomes. Advances inboth software and hardware are making it increasinglyviable to use the technology on mobile devicesand smart sensors.22 On the software side, severaltechnology vendors are creating compact machinelearning models that often require relatively littlememory but can still handle tasks such as imagerecognition and language translation on mobiledevices.23</OtherInformation></Objective></Goal><Goal><Name>Algorithmic Risks</Name><Description>Manage algorithmic risks.</Description><Identifier>_6ffd366c-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>4</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Nancy Albinson</Name><Description>Co-Author -- Nancy Albinson is a managing director with Deloitte &amp; Touche LLP and leader of Deloitte Risk &amp;Financial Advisory’s innovation program. She is based in Parsippany, NJ.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Dilip Krishna</Name><Description>Co-Author -- Dilip Krishna is the chief technology officer and a managing director with the Regulatory &amp;Operational Risk practice at Deloitte &amp; Touche LLP. He is based in New York City.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Yang Chu</Name><Description>Co-Author -- Yang Chu is a senior manager at Deloitte &amp; Touche LLP. She is based in San Francisco, CA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>William D. Eggers</Name><Description>Co-Author -- William D. Eggers is the executive director of Deloitte’s Center for Government Insights, where he isresponsible for the firm’s public sector thought leadership. He is based in Arlington, VA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Adira Levine</Name><Description>Co-Author -- Adira Levine is a strategy consultant at Deloitte Consulting LLP, where her work is primarily alignedto the public sector. She is based in Arlington, VA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>Chief data officers (CDOs), as the leaders of theirorganization’s data function, have an important roleto play in helping governments harness this newcapability while keeping the accompanying risks atbay.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Governments</Name><Description>THE PROBLEM OF ALGORITHMIC BIAS -- Governments have used algorithms to make various decisions in criminal justice, human services, healthcare, and other fields. In theory, this should lead to unbiased and fair decisions. However, algorithmshave at times been found to contain inherent biases, often as a result of the data used to train thealgorithmic model.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>For government agencies, the problem of biased input data constitutes one of thebiggest risks they face when using machine learning.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Court Systems</Name><Description>While algorithmic bias can involve a number of factors other than race, allegations of racial bias haveraised concerns about certain government applications of AI, particularly in the realm of criminaljustice. Some court systems across the country have begun using algorithms to perform criminal riskassessments, an evaluation of the future criminal risk potential of criminal defendants.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Judges</Name><Description>In nine USstates, judges use the risk scores produced in these assessments as a factor in criminal sentencing.However, criminal risk scores have raised concerns over potential algorithmic bias and led to calls forgreater examination.5</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>ProPublica</Name><Description>In 2016, ProPublica conducted a statistical analysis of algorithm-based criminal risk assessments inBroward County, Florida. Controlling for defendant criminal history, gender, and age, the researchersconcluded that black defendants were 77 percent more likely than others to be labeled at higher riskof committing a violent crime in the future.6 While the company that developed the tool denied thepresence of bias, few of the criminal risk assessment tools used across the United States have undergoneextensive, independent study and review.7</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Allegheny County, Pennsylvania</Name><Description>THE ALLEGHENY COUNTY APPROACH -- Some governments have begun building transparency considerations into their use of algorithms andmachine learning. Allegheny County, Pennsylvania provides one such example. In August 2016, thecounty implemented an algorithm-based tool—the Allegheny Family Screening Tool—to assess risks tochildren in suspected abuse or endangerment cases.13 The tool conducts a statistical analysis of morethan 100 variables in order to assign a risk score of 1 to 20 to each incoming call reporting suspectedchild mistreatment.14 Call screeners at the Office of Children, Youth, and Families consult the algorithm’srisk assessment to help determine which cases to investigate. Studies suggest that the tool has enableda double-digit reduction in the percentage of low-risk cases proposed for review as well as a smallerincrease in the percentage of high-risk calls marked for investigation.15Like other risk assessment tools, the Allegheny Family Screening Tool has received criticism for potentialinaccuracies or bias stemming from its underlying data and proxies. These concerns underscore theimportance of the continued evolution of these tools. Yet the Allegheny County case also exemplifiespotential practices to increase transparency. Developed by academics in the fields of social welfare anddata analytics, the tool is county-owned and was implemented following an independent ethics review.16County administrators discuss the tool in public sessions, and call screeners use it only to decide whichcalls to investigate rather than as a basis for more drastic measures. The county’s steps demonstrate oneway that government agencies can help increase accountability around their use of algorithms.</Description></Stakeholder><OtherInformation>How CDOs can manage algorithmic risks -- THE RISE OF advanced data analytics and cognitivetechnologies has led to an explosion inthe use of complex algorithms across a widerange of industries and business functions, as wellas in government. Whether deployed to predictpotential crime hotspots or detect fraud and abusein entitlement programs, these continually evolvingsets of rules for automated or semi-automateddecision-making can give government agencies newways to achieve goals, accelerate performance, andincrease effectiveness.However, algorithm-based tools—such asmachine learning applications of artificial intelligence(AI)—also carry a potential downside. Evenas many decisions enabled by algorithms have anincreasingly profound impact, growing complexitycan turn those algorithms into inscrutable blackboxes. Although often enshrouded in an aura ofobjectivity and infallibility, algorithms can bevulnerable to a wide variety of risks, including accidentalor intentional biases, errors, and fraud.</OtherInformation><Objective><Name>Risks</Name><Description>Understand the risks.</Description><Identifier>_6ffd3914-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>4.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Federal Agencies</Name><Description>Federal, state,and local governments are harnessing AI to solvechallenges and expedite processes—ranging fromanswering citizenship questions through virtualassistants at the Department of Homeland Securityto, in other instances, evaluating battlefield woundswith machine learning-based monitors.1</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>State Governments</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Local Governments</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Department of Homeland Security</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Smart Cities</Name><Description>In the coming years, machine learning algorithms willalso likely power countless new Internet of Things(IoT) applications in smart cities and smart militarybases.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Smart Military Bases</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Human Beings</Name><Description>While such change can be considered transformativeand impressive, instances of algorithmsgoing wrong have also increased, typically stemmingfrom human biases, technical flaws, usageerrors, or security vulnerabilities. For instance:</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Social Media</Name><Description>• Social media algorithms have come underscrutiny for the way they may influencepublic opinion.2</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Britain</Name><Description>• During the 2016 Brexit referendum, algorithmsreceived blame for the flash-crash of the Britishpound by six percent in two minutes.3</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Criminal Justice Systems</Name><Description>• Investigations have found that an algorithmused by criminal justice systems across theUnited States to predict recidivism rates isbiased against certain racial groups.4</Description></Stakeholder><OtherInformation>Understanding the risks -- Governments increasingly rely on data-driveninsights powered by algorithms...Typically, machine learning algorithms arefirst programmed and then trained using existingsample data. Once training concludes, algorithmscan analyze new data, providing outputs basedon what they learned during training and potentiallyany other data they’ve analyzed since. Whenit comes to algorithmic risks, three stages of thatprocess can be especially vulnerable:• Data input: Problems can include biases in thedata used for training the algorithm (see sidebar“The problem of algorithmic bias”). Other problemscan arise from incomplete, outdated, orirrelevant input data; insufficiently large anddiverse sample sizes; inappropriate data collectiontechniques; or a mismatch between trainingdata and actual input.• Algorithm design: Algorithms can incorporatebiased logic, flawed assumptions orjudgments, structural inequities, inappropriatemodeling techniques, or coding errors.• Output decisions: Users can interpretalgorithmic output incorrectly, apply it inappropriately,or disregard its underlying assumptions. ^The immediate fallout from algorithmic riskscan include inappropriate or even illegal decisions.And due to the speed at which algorithms operate,the consequences can quickly get out of hand. Thepotential long-term implications for governmentagencies include reputational, operational, technological,policy, and legal risks.</OtherInformation></Objective><Objective><Name>Approaches</Name><Description>Develop and adopt new approaches.</Description><Identifier>_6ffd3bc6-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>4.2</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name/><Description/></Stakeholder><OtherInformation>Taking the reins -- To effectively manage algorithmic risks, traditionalrisk management frameworks should bemodernized. Government CDOs should developand adopt new approaches that are built on strongfoundations of enterprise risk management andaligned with leading practices and regulatory requirements.Figure 1 depicts such an approach andits specific elements.</OtherInformation></Objective><Objective><Name>Strategy</Name><Description>Create an algorithmic risk management strategy and governance structure to manage technical and cultural risks.</Description><Identifier>_6ffd4094-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>4.2.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>European Union</Name><Description>InMay 2018, the European Union began enforcinglaws that require companies to be able to explainhow their algorithms operate and reach decisions.8</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>New York City Council</Name><Description>Meanwhile, in December 2017, the New York City Council passed a law establishing an AutomatedDecision Systems Task Force tostudy the city’s use of algorithmic systemsand provide recommendations. The bodyaims to provide guidance on increasingthe transparency of algorithms affectingcitizens and addressing suspected algorithmicbias.9</Description></Stakeholder><OtherInformation>STRATEGY, POLICY, AND GOVERNANCE -- Create an algorithmic risk management strategyand governance structure to manage technical andcultural risks. This should include principles, ethics,policies, and standards; roles and responsibilities;control processes and procedures; and appropriatepersonnel selection and training. Providing transparencyand processes to handle inquiries can alsohelp organizations use algorithms responsibly.From a policy perspective, the idea that automateddecisions should be “explainable” to thoseaffected has recently gained prominence, althoughthis is still a technically challenging proposition.</OtherInformation></Objective><Objective><Name>Life Cycle</Name><Description>Address potential issues in the algorithmic life cycle. </Description><Identifier>_6ffd435a-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>4.2.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Researchers</Name><Description>Researchershave developed a number of techniques to constructalgorithmic models in ways in which they canbetter explain themselves. One method involvescreating generative adversarial networks (GANs),which set up a competing relationship between twoalgorithms within a machine learning model. Insuch models, one algorithm develops new data andthe other assesses it, helping to determine whetherthe former operates as it should.10Another technique incorporates more directrelationships between certain variables into thealgorithmic model to help avoid the emergence ofa black box problem. Adding a monotonic layer toa model—in which changing one variable producesa predictable, quantifiable change in another—canincrease clarity into the inner workings of complexalgorithms.11</Description></Stakeholder><OtherInformation>DESIGN, DEVELOPMENT, DEPLOYMENT, AND USE -- Develop processes andapproaches aligned with the organization’salgorithmic risk managementgovernance structure to address potentialissues in the algorithmic life cycle fromdata selection, to algorithm design, tointegration, to actual live use in production.This stage offers opportunities to build algorithmsin a way that satisfies the growing emphasison “explainability” mentioned earlier.</OtherInformation></Objective><Objective><Name>Monitoring &amp; Testing</Name><Description>Establish processes for assessing and overseeing algorithm data inputs, workings, and outputs.</Description><Identifier>_6ffd458a-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>4.2.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Evaluators</Name><Description>Evaluators can not only assessmodel outcomes and impacts on alarge scale, but also probe how specificfactors affect a model’s individual outputs.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Researchers</Name><Description>For instance, researchers can examinespecific areas of a model, methodicallyand automatically testing differentcombinations of inputs—such as by inserting or removingdifferent parts of a phrase in turn—to helpidentify how various factors in the model affectoutputs.12</Description></Stakeholder><OtherInformation>MONITORING AND TESTING -- Establish processes for assessing and overseeingalgorithm data inputs, workings, and outputs,leveraging state-of-the-art tools as they becomeavailable. Seek objective reviews of algorithms byinternal and external parties.</OtherInformation></Objective><Objective><Name>Readiness</Name><Description>Ask important questions about your agency’s preparedness to manage algorithmic risks.</Description><Identifier>_6ffd4a58-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>4.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>Adopting effective algorithmic risk managementpractices is not a journey that government agenciesneed to take alone.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Researchers</Name><Description>The growing awareness ofalgorithmic risks among researchers, consumeradvocacy groups, lawmakers, regulators, and otherstakeholders should contribute to a growing body of knowledge about algorithmic risks and, over time,risk management standards. In the meantime, it’simportant for CDOs to evaluate their use of algorithmsin high-risk and high-impact situations andimplement leading practices to manage those risksintelligently so that their organizations can harnessalgorithms to enhance public value.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Consumer Advocacy Groups</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Lawmakers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Regulators</Name><Description/></Stakeholder><OtherInformation>Are you ready to managealgorithmic risks?A good starting point for implementing an algorithmicrisk management framework is to askimportant questions about your agency’s preparednessto manage algorithmic risks. For example:• Where are algorithms deployed in your governmentorganization or body, and how arethey used?• What is the potential impact should those algorithmsfunction improperly?• How well does senior management within yourorganization understand the need to managealgorithmic risks?• What is the governance structure for overseeingthe risks emanating from algorithms? ^The rapid proliferation of powerful algorithmsin many facets of government operations is in fullswing and will likely continue unabated for years tocome. The use of intelligent algorithms offers a widerange of potential benefits to governments, includingimproved decision-making, strategic planning, operationalefficiency, and even risk management. Butin order to realize these benefits, organizations willlikely need to recognize and manage the inherentrisks associated with the design, implementation,and use of algorithms—risks that could increaseunless governments invest thoughtfully in algorithmicrisk management capabilities.</OtherInformation></Objective></Goal><Goal><Name>DATA Act</Name><Description>Implement the DATA Act.</Description><Identifier>_6ffd4d32-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Dave Mader</Name><Description>Co-Author -- Dave Mader is the chief strategy officer for the civilian sector within Deloitte Consulting LLP’s Federalpractice. He is based in Arlington, VA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Tasha Austin</Name><Description>Co-Author -- Tasha Austin, a senior manager in Deloitte &amp; Touche LLP’s Federal practice, leads the DATA Actoffering for Deloitte’s Risk and Financial Advisory practice. She is based in Arlington, VA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Christina Canavan</Name><Description>Co-Author -- Christina Canavan, a managing director in Deloitte &amp; Touche LLP’s Federal practice, leads theAdvisory Analytics practice for Financial Risk Transactions and Restructuring. She is based inArlington, VA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>United States Postal Service</Name><Description>Having access to key facts can drive impressiveimprovements: When the United States PostalService compiled and standardized a number of itsdata sets, the office of the USPS Inspector General’sdata-modeling team was able to use them to identifyabout $100 million in savings opportunities, as wellas recover more than $20 million in funds lost topossible fraud.1</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>For government chief data officers (CDOs), oneof the key drivers for data transparency is the federalgovernment’s effort to implement wide-scale datainteroperability through the Data Accountabilityand Transparency Act of 2014 (DATA Act), whichseeks to create an open data set for all federalspending. If successful, the DATA Act could dramaticallyincrease internal efficiency and externaltransparency.2 However, our interviews with morethan 20 DATA Act stakeholders revealed some potentialchallenges to its implementation that couldbe important to address.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Officials</Name><Description>The DATA Act’s intentBefore addressing these implementation challenges,it may help to know how the DATA Act setsout to make information on federal expendituresmore easily accessible and transparent.Implementation of the DATA Act is still inits early stages; the first open-spending data setwent live in May 2017.3 If the act is successfullyimplemented, by 2022, spending data will flowautomatically from agency originators to interestedgovernment officials and private citizens throughpublicly available websites. This could save time andincrease efficiency across the federal government inseveral ways, possibly including the following:Spending reports would populate automatically.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Agency leaders</Name><Description>Agency leaders wouldn’t need to requestdistinct spending reports from different units oftheir agencies—the information would compileautomatically. For example, a user could see theDepartment of Homeland Security’s spending at asummary level or review spending at the componentlevel.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Congress</Name><Description>Congress could make appropriationsmore transparent. When crafting legislation,Congress could evaluate the impact of spendingbills with greater ease. Shifting a few sliders on adashboard could show the impact of proposedchanges to each agency’s budget. Negotiationscould be conducted using easy-to-digest pie chartsreflecting each proposal’s impact.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Auditors</Name><Description>Auditors would need to do less detectivework. Auditors would have direct access to datadescribing spending at a granular level. Rather thanoften digging through disparate records and unconnectedsystems, auditors could see an integratedmoney flow. Using data analytics, auditors couldgauge the cost-effectiveness of spending decisionsor compare similar endeavors in different agenciesor regions. These efforts could help root out fraud.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Citizens</Name><Description>Citizens could see where the moneygoes. With greater spending transparency, citizenscould have real-time clarity into how governmentdecisions might influence local grant recipients,nonprofits, and infrastructure. It could be as easyfor a citizen to see the path of every penny as itwould for an agency head.</Description></Stakeholder><OtherInformation>Implementing the DATA Act for greater transparency and accessibility -- With data an often-underutilized asset in the public sector, enhancing availabilityand transparency can make a big difference in enabling agencies to usedata analytics to their advantage—and the public’s. -- THOUGHTFUL USE OF data-driven insights canhelp agencies monitor performance, evaluateresults, and make evidence-based decisions.</OtherInformation><Objective><Name>Schema</Name><Description>Maintain a unified data format, or “schema,” to organize federal spending reports.</Description><Identifier>_6ffd4f76-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Office of Management and Budget</Name><Description>The DATA Act mandates that the White HouseOffice of Management and Budget (OMB) maintaina unified data format, or “schema,” to organize allfederal spending reports. This schema, known asDAIMS (DATA Act Information Model Schema),represents an agreement on how OMB and the Departmentof the Treasury want to categorize federalspending.5 It’s a common taxonomy that all agenciescan use to organize information, and it couldshape how the federal government approaches budgetingfor years to come. To allow other agenciesto connect to DAIMS, OMB has built open-sourcesoftware—the “Data Broker”—to help agenciesreport their data.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Department of the Treasury</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>State Governments</Name><Description>While the DATA Act deals with federal governmentdata, it can indirectly affect how state andlocal governments manage their data as well. Dataofficers from state and local governments willlikely need to be familiar with DAIMS and the DataBroker if they hope to collect grants from the federalgovernment.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Local Governments</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Contractors</Name><Description>And when contractors adopt federal protocols, they’ll likely prefer to report to states in a similar format.</Description></Stakeholder><OtherInformation>OMB’s data schema: The foundation for change -- The DATA Act has the potential to transformvarious federal management practices. While muchwork remains to be done, the technology to supportthe DATA Act has already been developed, givingthe act a strong foundation.4</OtherInformation></Objective><Objective><Name>Implementation</Name><Description>Address implementation challenges and approaches.</Description><Identifier>_6ffd5502-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.2</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>Federal Government</Name><Description>Thefederal government currentlyidentifies grant recipients and contractors usingDUNS, the Data Universal Numbering System, aproprietary system of identification numbers withnumerous licensing restrictions.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Data Universal Numbering System</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>State Partners</Name><Description>A transparentfederal data set won’t be able to incorporate newdata sets from state and local partners unlessthose partners also spend scarce resources on theDUNS system to achieve compatibility.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Local Partners</Name><Description/></Stakeholder><OtherInformation>As federal CDOs transform their organizationsto meet the DATA Act’s new transparency standards,they could face a number of challenges, bothcultural and technical.If users see the DATA Act as a reporting requirementrather than as a tool, they are unlikely to unlockits full potential. Bare minimum data sets, lackingin detail, might satisfy reporting requirements, butthey would fail to support effective data analytics.Likewise, users unfamiliar with the DAIMS systemmay never bother to become adept with it.Technical challengesalso threaten DATA Actimplementation. Legacyreporting systems may notbe compatible with DAIMS...Lastly, theDAIMS schema, while a monumental achievement,will continue to need improvement. The currentDAIMS schema fails to account for the full federalbudgeting life cycle. Therefore, the ability to use thedata to organize operations is incomplete at best.6With care and commitment, however, theseproblems can be surmountable. Two steps CDOscan take are: </OtherInformation></Objective><Objective><Name>Tool</Name><Description>Approach the DATA Act as a managerial tool.</Description><Identifier>_6ffd580e-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.2.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Managers</Name><Description>If managers use the DAIMS system to run their ownorganizations, the data they provide would be granularand more accurate.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Congress</Name><Description>That said, one of the bestways to convince managers to adopt DAIMS fordaily use will likely be through active congressionalbuy-in. If congressional budgeters and appropriatorsbegin relying on DAIMS-powered dashboardsto allocate funds, agency managers could naturallygravitate to the same data for budget submissions—and, eventually, for other management activities.</Description></Stakeholder><OtherInformation>Convince managers to see the DATA Actas a tool, not a chore. To truly fulfill the DATAAct’s promise, workplaces should approach it as amanagerial tool, not merely a reporting requirement.</OtherInformation></Objective><Objective><Name>Education &amp; Demonstration</Name><Description>Educate users and managers to show them the benefits. </Description><Identifier>_6ffd5a84-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.2.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Managers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>DAIMS Users</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Small Business Association</Name><Description>One of the test cases for Data Broker,the Small Business Association (SBA), worked withtechnology specialists on the federal government’s18F team to find uses for the new data system. Inthe process, they found mislabeled data, madeseveral data quality improvements, and evendiscovered discretionary funds that they hadthought were already committed.7 Agencieslike the SBA, which experienced significantimprovements, could evangelize the benefitsof clean, transparent data for decision-making tothe larger public sector community.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>18F Team</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Managers</Name><Description>Further, morecan be done to invest in the upskilling of managers.This could help managers to develop a vision forhow data can be used and begin to provide the resourcesneeded to get there.</Description></Stakeholder><OtherInformation>Education can encourageagencies to incorporate DAIMS data into their ownoperations.</OtherInformation></Objective><Objective><Name>Execution</Name><Description>Improve execution.</Description><Identifier>_6ffd5f98-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>For all its laudable intent, the DATA Act mayfail to deliver its full potential unless it is effectivelyexecuted. Some steps for the federal government toconsider include:</OtherInformation></Objective><Objective><Name>Governance</Name><Description>Establish a permanent governance structure.</Description><Identifier>_6ffd6344-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.3.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Currently, OMB and Treasury are responsiblefor managing data standards for spending data.While this fulfills the basic mandates of the DATAAct, experts acknowledge that, with their currentresources, these two agencies can’t do the workindefinitely.8 To ensure DAIMS’s flexibility and stability,a permanent management structure shouldoversee it for the long term.</OtherInformation></Objective><Objective><Name>Sources</Name><Description>Extract information directly from source systems.</Description><Identifier>_6ffd65ec-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.3.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Currently, when a government agencyawards a contract, it reports the contract data usingseveral old reporting systems, many of which havewell-documented accuracy problems.9 Currently,DAIMS extracts financial information from theseinconsistent sources. The first major revision toDAIMS should require agencies to extract contractinformation directly from their source awardsystems. Going straight to the source for bothfinancial and award data should lead to more efficient processing, boost data quality, and couldsave agencies time and effort.</OtherInformation></Objective><Objective><Name>Numbering System</Name><Description>Adopt a numbering system that anyone can use.</Description><Identifier>_6ffd6b1e-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.3.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Everyone, from local governments toAmerican businesses, should be encouraged to integratetheir own budgeting data with the federalgovernment’s. Instead of using a proprietarynumbering system that excludes participants, thegovernment could consider adopting an opensourceor freely available numbering system.</OtherInformation></Objective><Objective><Name>Budget Life Cycle</Name><Description>Expand the DAIMS to reflect the full budget life cycle. </Description><Identifier>_6ffd6f6a-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>5.3.4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>The federal budget follows a lifecycle, from the president’s proposed budget to congressionalappropriations to payments. To properlytrack the flow of funds through this life cycle, thespending data in DAIMS should reflect the budgetas something that evolves over time from the beginning,with the receipt of tax revenues to finalpayments to grantees and contractors.CDOs will likely recognize both the potentialbenefits of enhancing an organization’s ability toleverage data, and the challenges of changing theway public organizations manage data. CDOs wouldhave to thoughtfully manage through the barriersto realize the potential benefits of readily available,transparent data. Leaders would be wise to preparetheir own organizations for change even as theDATA Act takes hold at the federal level.</OtherInformation></Objective></Goal><Goal><Name>Health &amp; Science</Name><Description>Enable knowledge to flow freely across the scientific community.</Description><Identifier>_6ffd71f4-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>6</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Juergen Klenk</Name><Description>Co-Author -- </Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Melissa Majerol</Name><Description>Co-Author -- </Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>Open Science: In need of champions -- The health care sector is teeming with data.Electronic health records, technologies such assmart watches and mobile apps, and major advancesin scientific research—especially in the areasof imaging and genomic sequencing—have givenus volumes of medical and biological data over thelast decade. One might assume that such a data-richlandscape inherently accelerates scientific discoveries.However, reams of data alone cannot generatenew insights, especially when they exist in silos, asis often the case today.Open Science—the notion that scientific research,including data and research methodologies,should be open and accessible—can offer a solution.Without powerful champions, however, suchopenness may remain the exception rather than therule. Practicing Open Science inherently requirescross-sector collaboration as well as buy-in from thepublic. This is where government chief data officers(CDOs) could play a key role.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Scientific Community</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Dr. Jay Bradner</Name><Description>Consider cancer research. Dr. Jay Bradner, adoctor at a small Harvard-sponsored cancer lab,created a molecule called JQ1—a prototype for a drugto target a rare type of cancer. Rather than keepingthe prototype a secret until it was turned into anactive pharmaceutical substance and patented, thelab made the drug’s chemical identity available onits website for “open source drug discovery.” Theconcept of open source drug discovery borrowstwo principles from open source computing—collaborationand open access—and applied them topharmaceutical innovation.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Cancer Researchers</Name><Description>Scientists from aroundthe world were able to learn about the drug’s chemicalidentity so that they could experiment with it onvarious cancer cells. These scientists, in turn, havecreated new molecules to treat cancer that are beingtested in clinical trials.4 Collaborations like theseallow hundreds of minds to study the individualpieces of a complex problem, multiplying the usualpace of discovery.</Description></Stakeholder><OtherInformation>CDOs, health data, and the Open Science movement -- Now is the time for Open Science -- The early stages of the Open Science movementcan be traced back to the 17th century, when theidea arose that knowledge must flow freely acrossthe scientific community to enable and acceleratescientific breakthroughs that can benefit all ofsociety.1 Four centuries later, Open Science remainsan idea that has yet to be fully realized. However,collaborative tools and digital technologies aremaking the endeavor more achievable than everbefore. Rather than simply sharing knowledgein scientific journals, we now have the ability toshare electronic health records, patient-generateddata, insurance claims data—even genomic data—in standardized, interoperable formats through web-based tools and the cloud. Moreover, with advancedanalytics and cognitive technologies, we canprocess large volumes of data to identify complexpatterns that can lead to new discoveries in waysthat were almost unimaginable until recently. Usingthese data and tools is essential to achieving OpenScience’s so-called FAIR principles—that datashould be findable, accessible, interoperable, andreusable2 (see the sidebar, “What is FAIR?”).</OtherInformation><Objective><Name>Open Science</Name><Description>Accelerate Open Science.</Description><Identifier>_6ffd7726-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>6.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Federal Governments</Name><Description>Federal and state governments—and theirCDOs—have two unique levers that they can applyto encourage greater openness and collaboration:They hold enormous quantities of health data, andthey have the ability to influence policy and practice.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>State Governments</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Public Health Care Programs</Name><Description>US government health data derives from publicprograms like Medicare and Medicaid, which collectivelycover one in three people in the UnitedStates;5 government-sponsored disease registries;the Million Veteran Program (MVP), one of theworld’s largest medical databases, which has collectedblood samples and health information froma million veteran volunteers; and the NationalInstitutes of Health’s (NIH’s) recent All of Us initiative,a historic effort to gather data from 1 millionor more US residents to accelerate research andimprove health.6</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>National Institutes of Health</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Department of Health and Human Services</Name><Description>In addition, federal agencies suchas the Department of Health and Human Services(HHS), as well as a handful of states, cities, andcounties around the country, have begun hiringCDOs to help determine how data is collected, organized,accessed, and analyzed.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Cities</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Counties</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Project Open Data</Name><Description>According to ProjectOpen Data, an online public repository created byPresident Barack Obama’s Open Data Policy andExecutive Order,7 the CDO’s role is “part data strategistand adviser, part steward for improving data quality, part evangelist for data sharing, part technologist,and part developer of new data products.”8</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>CDOs looking to advance Open Science shouldconsider ways to meaningfully share more governmenthealth data and to encourage nongovernmentstakeholders, including academic researchers,health providers, and ordinary citizens, to participatein Open Science data platforms and sharetheir own data. To do so, they will need to addressthe various technological, policy, and culturalchallenges.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Academic Researchers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Health Providers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Citizens</Name><Description/></Stakeholder><OtherInformation>Government CDOs can help accelerate Open Science</OtherInformation></Objective><Objective><Name>Barriers</Name><Description>Overcome the barriers.</Description><Identifier>_6ffd7a6e-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>6.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Health Care Stakeholders</Name><Description>Looking ahead -- The proliferation of digital health data, coupledwith advanced computational capacity and interoperableplatforms such as Data Commons, givessociety the basic tools to practice Open Sciencein health care research. However, making OpenScience a reality will require all health care stakeholders,including ordinary citizens, to participate.13</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government CDOs</Name><Description>Government CDOs can accelerate the spread ofOpen Science in several ways. They can establishpolicies and governance principles that encouragedata-sharing. They can conduct education, outreach,and community engagement efforts to helpstakeholders understand why and how to share dataand to encourage them to do so. And they can serveas role models by making their own agencies’ dataavailable for appropriate public use.Like all important movements, Open Sciencewill likely face ongoing challenges. Those at thehelm will need to balance the opportunities it provideswith the inherent risks, including those relatedto data privacy and security. Of all the stakeholdersin scientific discovery, government CDOs may beamong the best placed to help society sort throughthese opportunities and risks. As public servants,they have every incentive to embrace a leadershiprole in promoting Open Science for the commongood.</Description></Stakeholder><OtherInformation>Overcoming the barriers: Technology, policy, and culture</OtherInformation></Objective><Objective><Name>Technology</Name><Description>MOVE GOVERNMENT HEALTH DATA TO THE CLOUD</Description><Identifier>_6ffd7d16-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>6.2.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>NIH</Name><Description>In an effort to develop thisinfrastructure, the NIH has begun piloting a “DataCommons,” a virtual space where scientists canstore, access, and share biomedical data and tools.Here, researchers can utilize “digital objects ofbiomedical research” to solve difficult problems togetherand apply cognitive computing capabilities ina single cloud-based environment.9This platform embraces theFAIR principles, includingthe need to safeguard thedata it contains with secureauthentication and authorizationprocedures. The pilotis due to be completed in2020,10 after which lessonslearned are expected to be incorporatedinto a number of permanent, interoperable, sustainablyoperated Data Commons spaces.A Data Commons, however, is only as good asthe quality and quantity of the health data it contains.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Researchers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Health Agency CDOs</Name><Description>Government health agency CDOs can playan important role in increasing participation inData Commons by moving their agency’s data fromon-premise storage units to large-scale cloud platformsthat are interoperable with the NIH’s DataCommons, making it more accessible. Equally importantis to improve the quality of the shared data,which means putting it in formats that are findable,interoperable, and reusable—that is to say, makingit machine-actionable.</Description></Stakeholder><OtherInformation>TECHNOLOGY: MOVING GOVERNMENT HEALTH DATA TO THE CLOUD -- Open Science requires a technological infrastructurethat allows data to be securely shared,stored, and analyzed.</OtherInformation></Objective><Objective><Name>Policy</Name><Description>EDUCATE STAKEHOLDERS AND IMPLEMENT DATA-SHARING REGULATIONS</Description><Identifier>_6ffd8284-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>6.2.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government CDOs</Name><Description>Government CDOs have an opportunity toovercome such barriers to data-sharing througha combination of education, support structures,and appropriate policies and governance principles.CDOs could conduct educational outreachto academics, health care providers, and otherstakeholders to clarify data privacy laws such as theHealth Insurance Portability and Accountability Act(HIPAA) and the Health Information Technologyfor Economic and Clinical Health (HITECH) Act.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Health Care Stakeholders</Name><Description>The goal would be to help these stakeholders understandthat, rather than prohibiting data-sharing,these laws merely define parameters around whenand how to share data. Through written materials,videos, and live workshops, CDOs can clarifyregulatory requirements to encourage data-sharingamong health care stakeholders and individualswho are being asked to share their personal healthinformation.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>In addition to educating stakeholders, CDOscan prompt agencies to take advantage of certainpolicies that allow government agencies to requiredata-sharing.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>NIH Director</Name><Description>The 21st-Century Cures Act, for instance,gives the director of the NIH the authority torequire that data from NIH-supported research beopenly shared to accelerate the pace of biomedicalresearch and discovery.11</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Researchers</Name><Description>Such policies must becomplemented with appropriate benefits for researcherswho share their data—for instance, givingsuch researchers appropriate consideration for additionalgrants and/or naming them as co-authorson publications that use their data.</Description></Stakeholder><OtherInformation>POLICY: EDUCATING STAKEHOLDERS AND IMPLEMENTING DATA-SHARING REGULATIONS -- The legal and regulatory landscape surroundingwhat data can be shared, with whom, and for whatpurpose can be a source of confusion and cautionamong health care providers and institutions thatcollect or generate health data. The real and/orperceived ethical, civil, privacy, or criminal risksassociated with data-sharing have led many researchersand health care stakeholders to avoiddoing so entirely unless they feel it is essential. This“better safe than sorry” approach can impede highimpact,timely, and resource-efficient discoveryscience. Furthermore, in academia, a researcher’scareer advancement can depend on his or her abilityto attract grant funding, which in turn depends onhis or her ability to generate peer-reviewed publications.In this competitive environment, researchershave little incentive to collaborate with and sharetheir valuable data with their peers. On top of thesebarriers, the effort and cost associated with makingdata FAIR are significant.</OtherInformation></Objective><Objective><Name>Culture</Name><Description>ENGAGE THE BROADER COMMUNITY</Description><Identifier>_6ffd85ea-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>6.2.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Health Care Stakeholders</Name><Description>One way of engaging health care stakeholdersand scientists is by giving them access to appropriategovernment data and tools so that they can beginusing shared data and seeing its value for themselves.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Scientists</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Communities</Name><Description>Another way is to seek innovative solutionsto health and scientific challenges using communityengagement models such as code-a-thons, contests,and crowdsourcing.12</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>The Public</Name><Description>CDOs can also encourage thegeneral public to ensure that their data contributesto Open Science by educating them on how theycan—directly or through patient advocacy organizations—encourage researchers and clinicians toshare the data they collect.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Patient Advocacy Organizations</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Private Individuals</Name><Description>Lastly, with private individualsincreasingly generating large volumes ofvaluable health data through wearables and mobiledevices, CDOs can help such individuals understandhow they could best share this data with researchers.</Description></Stakeholder><OtherInformation>CULTURE: ENGAGE THE BROADER COMMUNITY -- Open Science requires cross-sector participationand engagement from government entities, healthcare stakeholders, researchers, and the public. Aspart of their efforts to evangelize data-sharing,CDOs should consider engaging the broader communityby stoking genuine interest and appreciationof the crucial role data-sharing plays in science andinnovation and the benefits every player can gainfrom it.</OtherInformation></Objective></Goal><Goal><Name>Ethics</Name><Description>Manage data ethics.</Description><Identifier>_6ffd8964-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>7</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Christopher Wilson</Name><Description>Author -- Christopher Wilson is a research fellow at the University of Oslo and the Beeck Center for Social Impact + Innovation at Georgetown University. He is based in Oslo, Norway.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Ethics Networks</Name><Description>NETWORKS AND RESOURCES FOR MANAGING DATA ETHICS -- There are several nonprofit, private-sector, and research-focused communities and events that can beuseful for government CDOs.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Responsible Data Eebsite</Name><Description>The Responsible Data website curates an active discussion list on a broadrange of technology ethics issues.14</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>International Association of Privacy Professionals</Name><Description>The International Association of Privacy Professionals (IAPP) managesa community list serve,15 and the conference on Fairness, Accountability, and Transparency in Machine Learning convenes academics and practitioners annually.16</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>UK Government</Name><Description>Past activities and consultations like the UK government’s public dialogue on the ethics of data ingovernment can also provide useful information,17 and has resulted in the adoption of a governmentwidedata ethics framework, which includes a workbook and guiding questions for addressing ethicalissues.18</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>International Organization for Standardization</Name><Description>Responding to the EU General Data Protection Regulation (GDPR) regulations, the InternationalOrganization for Standardization (ISO) has set up a new project committee to develop guidelines toembed privacy into the design stages of a product or service.19</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Center for Democracy and Technology</Name><Description>Several other useful tools and frameworks have been produced. The Center for Democracy andTechnology (CDT) has developed a Digital Decisions Tool to help ethical decision-making into the designof algorithms.20</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Utrecht Data School</Name><Description>The Utrecht Data School has developed a tool, data ethics decision aid (DEDA) that iscurrently being implemented by various municipalities in the Netherlands.21</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Michigan Department of Transportation</Name><Description>The Michigan Departmentof Transportation has produced a decision-support tool dealing with privacy concerns surroundingintelligent transportation systems,22 and </Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>IAPP</Name><Description>the IAPP provides a platform for managing digital consentprocesses.23</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Sunlight Foundation</Name><Description>The Sunlight Foundation has developed a set of tools to help city governments ensure thatopen data projects map community data needs.24</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Journalism Groups</Name><Description>Many organizations also offer trainings and capacity development, including the IAPP,25 journalism and</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>O’Reilly Group</Name><Description>nonprofit groups like the O’Reilly Group,26 and </Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>National Institute of Standards and Technology</Name><Description>the National Institute of Standards and Technology, whichoffers trainings on specific activities such as conducting privacy threat assessments.27</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>European Public Sector Information Platform</Name><Description>Several white papers and reports also offer a general overview of issues and approaches, including theEuropean Public Sector Information Platform’s report on ethical and responsible use of open governmentdata28 and </Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Tilburg University</Name><Description>Tilburg University’s report on operationalizing public sector data ethics.29This list is not exhaustive, but it does illustrate the breadth of available resources, and might providea useful starting point for learning more.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>MERL Tech Conference</Name><Description>Participants in the MERL Tech Conference on technology formonitoring, evaluation, research, and learning also maintain a hackpad with comparable networks andresources for managing data ethics in the international development sector.30</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>International Development Sector</Name><Description/></Stakeholder><OtherInformation>Managing data ethics: A process-based approach for CDOs -- THE PROLIFERATION OF strategies for leveragingdata in government is driven in largepart by a desire to enhance efficiency, buildcivic trust, and create social value. Increasingly,however, this potential is shadowed by a recognitionthat new technologies and data strategies alsoimply a novel set of risks. If not approached carefully,innovative approaches to leveraging data caneasily cause harm to individuals and communitiesand undermine trust in public institutions. Manyof these challenges are framed as familiar ethicalconcepts, but the novel dynamics through whichthese challenges manifest themselves are much lessfamiliar. They demand a deep and constant ethicalengagement that will be challenging for many chiefdata officers (CDOs).To manage these risks and the ethical obligationsthey imply, CDOs should work on developinginstitutional practices for continual learning and interactionwith external experts. A process-orientedapproach toward data ethics is well suited for dataenthusiasts with limited resources in the fast-changingworld of new technologies. Prioritizingflexibility over fixed solutions and collaborationover closed processes could lower the risk of ethicalguidelines and safeguards missing their mark byproviding false confidence or going out-of-date...Recommendations for CDOs: A process-focused response -- Ethically motivated CDOs could find themselvesin a uniquely challenging situation. The dynamicnature of data and technology means that it is nearlyimpossible to anticipate what kinds of resourcesand expertise will be needed to meet the ethicalchallenges posed by data-driven projects before oneactually engages deeply with them. Even if it werepossible to anticipate this, however, the limitationsimposed by most government institutions wouldmake it difficult to secure all the resources andexpertise necessary, and the fundamentally ambiguousnature of ethical dilemmas makes it difficult toprioritize data ethics management over daily work.Progressively assessing and meeting these challengesrequires a degree of flexibility that mightnot come naturally to all institutional contexts. Butthere are a few strategies that can help.</OtherInformation><Objective><Name>Processes</Name><Description>PRIORITIZE PROCESSES, NOT SOLUTIONS</Description><Identifier>_6ffd8f36-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>7.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>NYC Task Force</Name><Description>In some contexts, it might make sense to formalizeprocesses, creating bodies similar to the NYCtask force mandated to assess equity, fairness, andaccountability in how the city deploys algorithms. Inother contexts, it may make more sense to consideralternative formats like data ethics lunches or short30-minute brainstorming sessions immediately followingstanding meetings and try to get everybodyon the same page about this being an effort to buildand sustain meaningful trust between constituentsand government.</Description></Stakeholder><OtherInformation>Whenever possible, CDOs should establishflexible systems for assessing and engaging withthe ethical challenges that surround data-drivenprojects. Identifying a group of people within andacross teams that are ready to reflect on these issuesand are willing to be on standby for discussionscan greatly enhance the efficiency of discussions.Setting up open invitations at the milestones andinflection points for every project or activity thathas a data component (see sidebar, “Networks andresources for managing data ethics”) can facilitateconstant attention. Also, it allows the project teamto step back and explore ways to embed privacyprinciples in the early design stages. Keeping thesediscussions open and informal can help create thesense of dedication and flexibility often necessary totackle complex challenges in contexts with limitedresources. Keeping them regular can help instill aninstitutional culture of being thoughtful about dataethics...Flexibility can be key to making this kind ofengagement effective, but it’s also important to be prepared. For each project, consider identifying akey set of issues or groups that are worth extra attention,and prioritize getting more than two peopleinto discussion regularly. Group conversations canhelp surface creative solutions and different pointsof view, and having them early can help prevent unpleasantsurprises.</OtherInformation></Objective><Objective><Name>Experts</Name><Description>ENGAGE WITH EXPERTS</Description><Identifier>_6ffd92ce-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>7.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Research Communities</Name><Description>Research communities regularly publish relevant reportsand white papers. </Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Networks</Name><Description>Government networks discussthe pros and cons of different policy options.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Civil Society Networks</Name><Description>Civilsociety networks advance cutting-edge thinkingaround data ethics and sometimes provide directsupport to government actors.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Private Sector Organizations</Name><Description>Increasingly, privatesector organizations, funders, consultants, andtechnology-driven companies are also offering resources.Becoming familiar with these communities is afirst step; just subscribing to a few RSS feeds canprovide prompts every day, flagging issues that needattention and honing it to keep ethical challengesfrom slipping through the cracks.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Funders</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Consultants</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Technology-Driven Companies</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Experts</Name><Description>Cultivating relationshipswith experts and advocates can provideimportant resources during crises. Attending conferencesand events can provide a host of insightsand contacts in this area.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Advocates</Name><Description/></Stakeholder><OtherInformation>ENGAGE WITH EXPERTS OUTSIDE THE BOX -- A process-based approach to managing dataethics will only be effective if teams have the capacityto address the risks that are identified, andthis will rarely be the case in resource-strappedgovernment institutions. CDOs should invest incultivating a broad familiarity with discourses ondata ethics and responsible data and the expertsand communities that drive those discourses. Doingso can help build the capacity of the teams andstakeholders inside government and also supportinnovative approaches to solving specific ethicalchallenges through collaboration.Many sources of information and expertiseare available for managing data ethics. </OtherInformation></Objective><Objective><Name>Internal Processes</Name><Description>OPEN UP INTERNAL ETHICAL PROCESSES</Description><Identifier>_6ffd95c6-31cd-11ea-842c-c5da1d83ea00</Identifier><SequenceIndicator>7.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Open Source Digital Security Systems</Name><Description>Open source digital security systems provide anillustrative example.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Digital Security Experts</Name><Description>A number of services are availablefor encrypting communications, but digitalsecurity experts recommend using open sourcedigital security software because its source code isconsistently audited and reviewed by an army ofpassionate technologists who are vigilant to vulnerabilitiesor flaws. As it is not possible to audit closedsource encryption tools in the same way, it is notpossible to know when and to what degree the securityof those tools has been compromised.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Data Programs</Name><Description>In much the same way, government data programsmay be working to keep information orpersonal data secure and private, but by havingopen discussions about how to do so, they typicallybuild trust with the communities they are trying toserve. They also open up the possibility of input andcorrections that can improve data ethics strategiesin the long and short run.This kind of openness could involve describingprocesses in op-eds, blog posts, or event presentationsor inviting the occasional expert to data ethicslunches or the other flexible activities describedabove. Or it might involve the publication of documents,regular interaction with the press, or a morestructured way of engaging with the communitiesthat are likely to be affected by data ethics.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>Whateverthe mechanism or the particular constraints onCDOs, a default inclination toward open processeswill contribute toward building trust and creatingeffective data ethics strategies.</Description></Stakeholder><OtherInformation>Perhaps most importantly, process-focused approachesto managing data ethics should be openabout their processes. Though some governmentinformation will need to be kept private for securityreasons, CDOs should encourage discussionsabout keeping the management of ethics open andtransparent whenever possible. This adheres to animportant emerging norm regarding open government,but it’s also critical for making data ethicsstrategies effective.</OtherInformation></Objective></Goal><Goal><Name>Data Lakes</Name><Description>Maximize the data lake investment</Description><Identifier>_cbb659e0-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Paul Needleman</Name><Description>Co-Author -- Paul Needleman is a specialist master with Deloitte Consulting LLP. He is based in Rosslyn, Va.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Eric Rothschild</Name><Description>Co-Author -- Eric Rothschild is a senior consultant with Deloitte Consulting LLP. He is based in Rosslyn, Va.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Stephen Schiavone</Name><Description>Co-Author -- Stephen Schiavone is a business technology analyst with Deloitte Consulting LLP. He is based inRosslyn, Va.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Business Users</Name><Description>Data lakes are special,in part, because they provide business users withdirect access to raw data without significant ITinvolvement. This “self-service” access lets usersquickly analyze data for insights.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Enterprises</Name><Description>Because they storethe full spectrum of an enterprise’s data, data lakescan break down the challenge of data silos that oftenbedevil data users.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>IT Middlemen</Name><Description>Byenabling easy access to enterprise data, data lakesallow subject matter experts to perform data analyticswithout going through an IT “middleman.”</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Users</Name><Description>Atthe same time, however, these data lakes mustprovide users with enough context for the data to beusable—and useful.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Business Leaders</Name><Description>To invest or not to invest?The challenges associated with traditionaldata storage platforms have led today’sbusiness leaders to look for modern, forward-looking,flexible solutions.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>Data lakes areone such solution that can help governmentagencies utilize information in ways vastlydifferent than was previously possible.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>CDOs</Name><Description>Providing broad access to raw data presentsboth a challenge and an opportunity for CDOs...CDOs can play a major role in the developmentof a data lake, providing a strategic vision thatencourages usability, security, and operationalimpact...It would be easy to say that there is aone-size-fits-all approach and that everyorganization should have a data lake, but this isnot true. A data lake is not a silver bullet, and itis important for CDOs to evaluate their organization’sspecific needs before making that investment.By planning properly, understanding user needs,educating themselves on the potential pitfalls, andfostering collaboration, a CDO can gain a solidfoundation for making the decision.</Description></Stakeholder><OtherInformation>Pump your own data: Maximizing the data lake investment -- ORGANIZATIONS ARE CONSTANTLY lookingfor better ways to turn data into insights,which is why many government agenciesare now exploring the concept of data lakes.Data lakes combine distributed storage withrapid access to data, which can allow for fasteranalysis than more traditional methods such asenterprise data warehouses...Implemented correctly, datalakes provide insight at the point of action, and giveusers the ability to draw on any data at any time toinform decision-making.Data lakes store information in its raw andunfiltered form—whether it is structured, semi-structured,or unstructured. A data lake performslittle automated data cleansing or transformation.Instead, data lakes shift the responsibility of datapreparation to the business.</OtherInformation><Objective><Name>Data Swamps</Name><Description>Avoid a data swamp</Description><Identifier>_cbb662fa-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Establishing the data lake platform: Avoiding a data swamp -- A poorly executed data lake is known as a dataswamp: a place where data goes in, but does notcome out. To ensure that a data lake provides valueto an organization, a CDO should take some importantsteps.</OtherInformation></Objective><Objective><Name>Metadata</Name><Description>Help users make sense of data</Description><Identifier>_cbb66c14-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.1.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>METADATA: HELP USERS MAKE SENSE OF THE DATA -- Imagine being turned loose in a library withouta catalog or the Dewey Decimal System, and wherethe books are untitled. All the information in thebooks is there, but good luck turning it into usefulinsight. The same goes for data lakes: To reap thedata’s value, users need a metadata “map” to locate,make sense of, and draw relationships among theraw data stored within. This metadata layer providesadditional context for data that flows throughto the data lake, tagging information for ease of uselater on.Too often, raw data is stored with insufficientmetadata to give the user enough context to makegainful use of it. CDOs can help combat this situationby acting as a metadata champion. In this capacity,the CDO should make certain that the metadata inthe data lakes he or she oversees is well understoodand documented, and that the appropriate businessusers are aware of how to use it.</OtherInformation></Objective><Objective><Name>Security</Name><Description>Control access for individual roles</Description><Identifier>_cbb676be-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.1.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>SECURITY: CONTROL ACCESS FOR INDIVIDUAL ROLES -- By putting appropriate security and controlsin place, CDOs will be better positioned to meetincreasingly stringent compliance requirements.Given the vast amount of information data lakestypically contain, CDOs need to control which usershave access to which parts of the data.Role-based access control (RBAC) is a controlmechanism defined around roles and privilegesthrough security groups. The components of RBAC—such as role permissions, user roles, and role-to-rolerelationships—make it simple to grant individualsspecific access and use rights, minimizing the riskof noncleared users accessing sensitive data. Withinmost data lake environments, security typically canbe controlled with great precision, at a file, table,column, row, or search level.Besides improving security, role-based accesssimplifies the user experience because it providesusers with only the data they need. It also enhancesconsistency, which can build users’ trust in theaccessed data; this, in turn, can increase useradoption.</OtherInformation></Objective><Objective><Name>Data Preparation</Name><Description>Equip users to cleanse data sets</Description><Identifier>_cbb67c90-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.1.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Business Users</Name><Description/></Stakeholder><OtherInformation>DATA PREPARATION: EQUIP USERS TO CLEANSE DATA SETS -- Preparing a new data set can be an extremelytime-consuming activity that can stymie data analysisbefore it begins. To obtain a reliable analyticoutput, it’s usually necessary to cleanse, consolidate,and standardize the data going in—and with a datalake, the responsibility of preparing the data fallslargely into the hands of the business users. Thismeans the CDO must work with business users togive them tools for data prep.Thankfully, software is emerging to help withthe work of data preparation. The IT organizationshould work collaboratively with the datalake’s business users to create tools and processesthat allow them to prepare and customizedata sets without needing to know technicalcode—and without the IT department’s assistance.Equipped with the right tools and know-how,business data users can prepare data efficiently,allowing them to focus the bulk of their efforts ondata analysis.</OtherInformation></Objective><Objective><Name>Enablement</Name><Description>Allow users to use familiar tools</Description><Identifier>_cbb681d6-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.1.4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>ENABLEMENT: ALLOW USERS TO USE FAMILIAR TOOLS -- Self-service data analysis will go more smoothlyif users can use familiar tools rather than havingto learn new technologies. CDOs should strive toensure that the business’s data lake(s) will be compatiblewith the tools the business currently uses.This will greatly enhance the data lake platform’seffectiveness and adoption. Fortunately, data lakessupport many varieties of third-party softwarethat leverage SQL-like commands, as well as opensource languages such as Python and R.</OtherInformation></Objective><Objective><Name>Governance</Name><Description>Maintain controls for non-IT resources</Description><Identifier>_cbb68ab4-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.1.5</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>IT</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Business Users</Name><Description/></Stakeholder><OtherInformation>GOVERNANCE: MAINTAIN CONTROLS FOR NON-IT RESOURCES -- Once users have access to data, they will use it—which is the whole point of self-service. But what if auser makes an error in data extraction, leading to aninaccurate result? Self-service is fine for exploringdata, but for mission-critical decisions or widespreaddissemination, the analytical outcomes mustbe governed in a way to guarantee trust.One approach to maintaining appropriate governancecontrols is to use “zones” for data access andsharing, with different zones allowing for differentlevels of review and scrutiny (figure 1). This allowsusers to explore data for inquiry without exhaustivereview while simultaneously requiring that data thatwill be broadly shared or used in critical decisionswill be appropriately vetted. With such controls inplace, a data lake’s ecosystem can perform nimblywhile limiting the impact of mistakes in extractionor interpretation.Figure 1 illustrates one possible governancestructure for a data lake ecosystem in which differentzones offer appropriate governance controls:• Zone 1 is owned by IT and stores copies of theraw data through the ingestion process. Thiszone contains the least trustworthy data andrequires the most vetting.• Zone 2 is where business users can create theirown data sets based on raw data from zone 1 as well as external data sources. Zone 2 would betrusted for group (i.e., office or division) use, andcould be controlled by group-sharing settings.• Zone 3 data sets, maintained by IT, are vettedand stored in optimal formats before beingshared with the broader organization. Onlydata in zone 3 would be trusted for broadorganizational uses.</OtherInformation></Objective><Objective><Name>Empowerment</Name><Description>Empower people to use data lakes</Description><Identifier>_cbb68faa-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Employees</Name><Description/></Stakeholder><OtherInformation>Adopting a culture of self- service analytics: Empowering people to use data lakes -- Implementing a data lake is more than a technicalendeavor. Ideally, the establishment of a datalake will be accompanied by a culture shift thatembeds data-driven thinking across the enterprise,fostering collaboration and openness amongvarious stakeholders. The CDO’s leadership throughthis transition is critical in order to give employeesthe resources and knowledge needed to turn datainto action.</OtherInformation></Objective><Objective><Name>Leaders</Name><Description>Invest in data leaders</Description><Identifier>_cbb69504-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.2.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Leaders</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Senior Business Leaders</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Data Champions</Name><Description/></Stakeholder><OtherInformation>INVEST IN DATA LEADERS -- CDOs are responsible for more than just the datain the data lake; they are also responsible for helpingto equip the workforce with the data skills they needto effectively use the data lake. One way to helpachieve this is for CDOs to advocate for and investin employees that have the necessary skills, attitude,and enthusiasm. Specialized trainings, town halls,data boot camps—a variety of approaches may beneeded to foster not only the technical skills, butthe courage to change outdated approaches thattrap data in impenetrable silos. The best CDOs willcreate an organization of data leaders.CDOs may need to work with senior businessleaders and HR in the drive for change. They shouldstrive to overcome barriers, highlight data championsthroughout the organization, and lead byexample.</OtherInformation></Objective><Objective><Name>Governance</Name><Description>Practice nimble governance</Description><Identifier>_cbb69d56-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.2.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Governance Bodies</Name><Description/></Stakeholder><OtherInformation>PRACTICE NIMBLE GOVERNANCE -- Governance over data lakes needs to walk avery fine line to support effective gatekeeping forthe data lake while not impeding users’ speed orsuccess in using it. Traditionally, governance bodiesfor data defined terms, established calculations,and presented a single voice for data. While thisis still necessary for data lakes, governance bodiesfor a data lake also should establish best practicesfor working with the data. This includes activitiessuch as working with business users to review dataoutputs and prioritizing ingestion within the datalake environment. Organizations should establishthorough policy-based governance to control wholoads which data into the data lake and when or howit is loaded.</OtherInformation></Objective><Objective><Name>Technology</Name><Description>Keep current on technology</Description><Identifier>_cbb6a2ec-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>8.2.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>KEEP CURRENT ON THE TECHNOLOGY -- Technology is never static; it will always evolve,improve, and disrupt at a dizzying speed. The technologysurrounding data lakes is no exception. Thus, CDOs must continue to make strategic investmentsin their data lake platforms to update them withnew technologies.To do this effectively, CDOs must educate themselvesabout current opportunities for improvingthe data lake and about new technologies that willreduce users’ burden. Doing so will open up theability for more users to use data in their everydaywork and decisions. Keeping oneself up to date isstraightforward: Read journals and trades,attend conferences and meetups, talk to theusers, and be critical of easy-sounding solutions.This will empower a CDO to siftthrough the vaporware, buzzwords, andflash to identify tactical, practical, andnecessary improvements.</OtherInformation></Objective></Goal><Goal><Name>Data Strategy</Name><Description>Define and implement a data strategy</Description><Identifier>_cbb6aad0-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Max Duhe</Name><Description>Co-Author -- Max Duhe is a consultant with Deloitte Consulting LLP. He is based in Arlington, Va.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Matt Gracie</Name><Description>Co-Author -- Matt Gracie is a managing director with Deloitte Consulting LP’s strategy and analytics team. He isbased in Arlington, Va.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Chris Maroon</Name><Description>Co-Author -- Chris Maroon is a senior consultant with Deloitte Consulting LLP. He is based in Arlington, Va.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Tess Webre</Name><Description>Co-Author -- Tess Webre is a senior consultant with Deloitte Consulting LLP. She is based in Rosslyn, Va.</Description></Stakeholder><OtherInformation>Data as an asset: Defining and implementing a data strategy -- HOW DO YOU get your organization to value data? Accurate data is the fuel that propelsgovernment organizations toward achievingtheir mission. Increasingly, government organizationsagree it is time to view data as a criticalstrategic asset and treat it accordingly.1Managing and leveraging data typically falls tothe chief data officer (CDO). In the Big Data ExecutiveSurvey of 2017, 41.4 percent of the executivessurveyed believed that the CDO’s primary roleshould be to manage and leverage data as an enterprisebusiness asset.2 However, many governmentorganizations fail to invest in the resources necessaryto realize the data’s inherent value.It is easy to become overwhelmed by the challengeof turning data from an afterthought into acore facet of business operations. Organizations canbecome paralyzed because they don’t know whereto begin. But CDOs can take comfort in knowingthat change doesn’t happen overnight.To unlock the value of an organization’s data, aCDO should develop and implement a clear datastrategy. The data strategy can help organizationstake a strategic view of data and use it more effectivelyto drive results. The best data strategiesare generally tailored to the organization’s needsand help the CDO engage necessary stakeholders,plan for the future, implement strategic projects,develop partnerships across the organization, andemphasize successes to drive a strategic mindset.</OtherInformation><Objective><Name>Success</Name><Description>Define a data strategy with success in mind</Description><Identifier>_cbb6b3cc-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Where to begin: Defining a data strategy with success in mind -- A data strategy provides an organization withdirection. CDOs can use the data strategy to organizedisparate activities, consolidate siloed data,and orient the organization toward a cohesive and unified goal. The aim is to set the stage fortreating data as an asset, resulting in improved decision-making, enhanced user insights, and greatermission effectiveness.</OtherInformation></Objective><Objective><Name>Tailoring</Name><Description>Tailor data strategies to organizations' unique needs</Description><Identifier>_cbb6bbce-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.1.1</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>US Navy CIO</Name><Description>For instance, the US Navy CIO’s data strategyemphasizes data analytics and data managementto enhance combat capabilities.3</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Department of Health and Human Services</Name><Description>Similarly, theDepartment of Health and Human Services’ datastrategy focuses on consolidating data repositoriesto create a shareable data environment for all relevantstakeholders.4</Description></Stakeholder><OtherInformation>TAILORING A DATA STRATEGY TO AN ORGANIZATION’S UNIQUE NEEDS -- Every organization is different; there is no definitivechecklist for a data strategy. Successful datastrategies come in many shapes and sizes, tailoredto each organization’s strengths and weaknesses.CDOs who are unsure of theirorganization’s strengths and weaknessesbenefit from an assessment of their datamaturity. Assessments are intended to provide apulse check that CDOs can use to prioritize goalsand initiatives within the data strategy to meetthe organization’s unique needs. With this understandingof strengths and weaknesses, CDOs cantailor their strategy to build upon organizationaldata opportunities while being cognizant of limitations.The aims of an organization’s data strategyshould align with the overall mission and goals.</OtherInformation></Objective><Objective><Name>People</Name><Description>Consider the human side</Description><Identifier>_cbb6c18c-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.1.2</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Michael Flowers</Name><Description>For example, New York City’s first chief analyticsofficer, Michael Flowers, addressed several complexproblems through a data strategy that emphasizedengagement of all data owners across the local government.“Our big insight was that data trapped inindividual agencies should be liberated and usedas an enterprise asset,” he said.5 Flowers’ effortsled to the development of New York City’s dataintegration platform, which now allows differentparts of the local government to share data witheach other to collaborate and solve problems.6</Description></Stakeholder><OtherInformation>IT’S ALL ABOUT THE PEOPLE -- To be effective, a data strategy should alsoconsider the human side: owners, stakeholders,analysts, and other users. Organizations that encouragestaff to think about information and data asa strategic asset can extract more value from theirsystems...Gaining buy-in across the organizationis instrumental in developing a successfuldata strategy, as is understandingall relevant organizational needs. The CDOshould engage all parts of the organizationfrom day one. Without inputfrom key stakeholders, the CDO mayfail to incorporate critical organizationalconsiderations into the datastrategy.</OtherInformation></Objective><Objective><Name>Planning</Name><Description>Plan for the future</Description><Identifier>_cbb6ca74-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.1.3</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>City of San Francisco</Name><Description>As an example, the City of San Francisco addressesthis challenge by regularly revisiting itsdata strategy. The City reviews and revises its datastrategy each year—including refining its mission,vision, and approach—all while adhering to a setlist of core goals. This periodic review of its datastrategy keeps the City’s approach to data use upto date while sustaining accountability for pursuingthe City’s overall strategy. 7</Description></Stakeholder><OtherInformation>PLANNING FOR THE FUTURE -- Nothing is stationary. CDOs should recognizethat not only will their organization change, but sowill various industry tools and technologies, as wellas broader government policies and practices. It isimperative to plan and establish a data strategy thataccounts for future changes. A flexible data strategycan open up the ability for the organization to continueto use data as an asset for the long term.</OtherInformation></Objective><Objective><Name>Implementation</Name><Description>Turn the document into a movement</Description><Identifier>_cbb6cfb0-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>How to implement a data strategy: Turning a document into a movement -- Implementing a data strategy is a daunting task.One common difficulty is that many organizationsare hesitant to change legacy IT operations—especiallyfor government, whose budgeting process canmake even small changes difficult to implement.However, difficult does not equate to impossible.CDOs can nudge their organization toward alignmentwith the data strategy’s principles and goals.</OtherInformation></Objective><Objective><Name>Action</Name><Description>Transform the strategy into action</Description><Identifier>_cbb6d5fa-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.2.1</SequenceIndicator><Stakeholder StakeholderTypeType="Organization"><Name>US Air Force</Name><Description>For example,while developing the US Air Force’s data strategy,the CDO identified manpower shortages as a criticalissue. The CDO prioritized this limitation early onin the implementation of the data strategy and developeda proof of concept to address it.8</Description></Stakeholder><OtherInformation>TRANSFORMING STRATEGY INTO ACTION -- Even the most brilliant strategy will not improvean organization’s use of its data assets if it sits ona shelf. Converting a data strategy from a piece ofpaper into a state of mind can be messy, and CDOsshould be realistic about the pace of change, especiallyearly on. The most effective approach in theface of organizational inertia can be to set realisticexpectations and identify opportunities to showvalue early on.Once the data strategy is developed, CDOsshould identify and list key business issues that thedata strategy is designed to address or solve. Forexample, will the data strategy enable the organizationto meet upcoming regulatory or legislativedeadlines? Are there existing modernization effortsunderway that require a data conversion? Is there aparticular weakness from the assessment that canbe addressed by implementing a data governancecouncil? CDOs can develop a list of projects by identifyingspecific ways the data strategy can addressthese issues.It is important to prioritize issues that will addthe most value to the organization. To establish thedata strategy’s credibility and utility, it’s helpful tostart with high-visibility projects that draw on keycomponents of the data strategy and that supportthe CDO’s own key goals.What defines a good opportunity will be differentfor each organization. Finding the right project requiresthe CDO to have a clear understanding ofthe organization’s wants and needs.</OtherInformation></Objective><Objective><Name>Victory</Name><Description>Turn action into victory</Description><Identifier>_cbb6df5a-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.2.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Partnerships</Name><Description>The best partnerships are those that aremutually beneficial, where all parties are investedin the effort’s outcome. All parties should have skinin the game; this way, once the solution is deployed,everyone can declare victory.An effective partnership can be maintained bysimple, frequent, prioritized, and actionable communication.Simple and frequent communicationkeeps all parties informed about progress and minimizesnegative impressions from minor setbacks.Delivering prioritized and actionable informationenables all parties to act efficiently, and fosters ashared sense of ownership.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Project Teams</Name><Description>One effective partnership model can be for theCDO to share resources with partners across theorganization. For example, a member of the CDO’steam could work with a partner for a limited timeto implement a specific project. This could benefitboth teams: The project team gets an additional resource,and the CDO can be confident that the workaligns with the overall strategy. If a team membercannot be spared, the CDO can provide the projectteam with tools, subject-matter expertise, or otherassets. Not only does this increase the odds that theproject will meet its objectives, but it also affordsthe CDO greater control over the project’s alignmentwith the data strategy.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>US Department of Transportation</Name><Description>An example of such a partnership is the USDepartment of Transportation’s (DOT’s) developmentof a national transit map in collaborationwith several state and local transportation organizations.9 In this effort, the DOT provided technicalassistance to local transit agencies, who also benefitedfrom having their data made publicly available.</Description></Stakeholder><OtherInformation>TURNING ACTION INTO VICTORY -- A CDO can improve the chances of a project’ssuccess by developing partnerships across the organization.</OtherInformation></Objective><Objective><Name>Mindset</Name><Description>Turn victory into a strategic mindset</Description><Identifier>_cbb6e4b4-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.2.3</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>HHS CDO</Name><Description>For example, enhanced dataavailability and a strategic analysis enabled theCDO of the US Department of Health and HumanServices’ Office of the Inspector General to detecthundreds of millions of dollars in fraud in 2017.10CDOs can point to victories like these when makinga budgetary case for additional resources, technologies,or capabilities.</Description></Stakeholder><OtherInformation>TURNING A VICTORY INTO A STRATEGIC MINDSET -- Success breeds success, and CDOs should capitalizeon every victory...Additionally, a CDO can usevictories to foster excitement and buy-in from theorganization.Every victory counts. After one victory, find thenext. Find another project, turn it into a success,and publicize it appropriately. Engage the energeticparticipants, and continue to work on the less-thanenthusiasticones.</OtherInformation></Objective><Objective><Name>Beginning</Name><Description>Begin now</Description><Identifier>_cbb6ea9a-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>9.2.4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>THE FUTURE STARTS NOW -- The years to come will present many uniqueopportunities and challenges for data management.CDOs are uniquely positioned to guide organizationsthrough the process of managing data andunlocking its value.To be successful in this effort, a CDO must havea nuanced understanding of the organization’scurrent data culture, resources, and opportunitiesfor improvement. Through this understanding,CDOs can develop and implement an actionabledata strategy to achieve the desired future state.Understanding how to create and tailor a datastrategy will be a critical skill for CDOs. By carefullyselecting and implementing a data strategy andcapitalizing on victories, CDOs can position theirorganizations for success in making use of data as avaluable strategic asset.</OtherInformation></Objective></Goal><Goal><Name>Tokenization</Name><Description>Enable data-sharing without compromising privacy</Description><Identifier>_cbb6f3be-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10</SequenceIndicator><Stakeholder StakeholderTypeType="Person"><Name>Tab Warlitner</Name><Description>Co-Author -- Tab Warlitner, a principal with Deloitte Consulting LLP and a member of its board of directors, isthe lead client service partner supporting the state of New York and New York City. He is based inArlington, VA.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>John O’Leary</Name><Description>Co-Author -- John O’Leary, a senior manager with Deloitte Services LP, is the state and local government researchleader for the Deloitte Center for Government Insights. He is based in Boston.</Description></Stakeholder><Stakeholder StakeholderTypeType="Person"><Name>Sushumna Agarwal</Name><Description>Co-Author -- Sushumna Agarwal is a senior analyst with the Deloitte Center for Government Insights, DeloitteServices LP. She is based in Mumbai.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Department of Homeland Security</Name><Description>Fromthe 9/11 terrorist attacks to fatal failures of childprotective services, we are often left to wonder:What if government’s left hand knew everything ithad in its right hand?</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Child Protective Services</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>That’s a tough ideal to attain for many governmentagencies, both in the United States and around theworld. Today, much of the information held ingovernment programs is isolated in siloeddatabases, limiting the ability to mine the data forinsights.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Agency Lawyers</Name><Description>Attempts to share this data throughinteragency agreements tend to be clunky at bestand nightmarish at worst, with lawyers frommultiple agencies often disagreeing over themeaning of obscure privacy provisions written bydisparate legislative bodies. No fun at all.This isn’t because agencies are being obstructionist.Rather, they’re acting with the best of intentions:to protect privacy.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Programs</Name><Description>Most government programs—such as Supplemental Nutrition AssistanceProgram (SNAP), Medicare, UnemploymentInsurance, and others—have privacy protectionsbaked into their enabling legislation. Limitingdata-sharing among agencies is one way tosafeguard citizens’ sensitive data against exposureor misuse. The fewer people have access to the data,after all, the less likely it is to be abused.The flip side, though, is that keeping the dataseparate can compromise agencies’ ability toextract insights from that data. Whether one isapplying modern data analytics techniques or justeyeballing the numbers, it’s usually best to workwith a complete view of the data, or at least themost complete view available. That can be hardwhen rules governing data-sharing preventagencies from combining their individual datapoints into a complete picture.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Supplemental Nutrition Assistance Program</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Medicare</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Unemployment Insurance Program</Name><Description/></Stakeholder><OtherInformation>Data tokenization for government: Enabling data-sharing without compromisingprivacy -- WE’VE ALL HEARD THE stories: If onlyinformation housed in one part ofgovernment had been available toanother, tragedy might have been averted...What if data could be shared across agencies,without compromising privacy, in a way that couldenable the sorts of insights now possible throughdata analytics?That’s the promise—or the potential—ofdata tokenization.</OtherInformation><Objective><Name>Sensitive Data</Name><Description>Replace sensitive data with substitute characters</Description><Identifier>_cbb6f936-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>How data tokenization works -- Data tokenization replaces sensitive data withsubstitute characters in a manner similar to datamasking or redaction. Unlike with the latter twoapproaches, however, the sender of tokenized dataretains a file that matches the real and tokenizeddata. This “token,” or key file, does two things.First, it makes data tokenization reversible, so thatany analysis conducted on the tokenized data canbe reconstituted by the original agencies—withadditional insights gleaned from other data sources.Second, it makes the tokenized data that leaves theagency virtually worthless to hackers, since it isdevoid of identifiable information.A simplified example can help illustrate howtokenized data can allow for personalized insightswithout compromising privacy. Imagine that youare a child support agency with the followinginformation about an individual:Name: Marvin BealsDate of birth: September 20, 1965Street address: 23 Airway DriveCity, state, zip: Lewiston, Idaho, 83501Gender: MaleHighest education level: Four-year collegeYou might tokenize this data in a way that keepscertain elements “real” (gender and education level,for example), broadens others (such as bytokenizing the day and month of birth but keepingthe real year, or tokenizing the street address butkeeping the actual city and state), and fullytokenizes still other elements (such as theindividual’s name). The result might looksomething like this:Name: Joe ProustDate of birth: May 1, 1965Street address: 4 Linden StreetCity, state, zip: Lewiston, Idaho, 83501Gender: MaleHighest education level: Four-year collegeYou could readily share this tokenized data with athird party, as it isn’t personally identifiable. Butyou could also combine this tokenized data with,for example, bank data that can predict what a54-year-old male living in that zip code is likely toearn, how likely he is to repay loans, and so forth.Or you could combine it with similarly tokenizeddata from a public assistance agency to learn howlikely a male of that age in that geographical area isto be on public assistance. After analysis, you (andyou alone!) could reverse the tokenization processto estimate—with much greater accuracy—howlikely Marvin is to be able to pay his child support.Going deeper, you could work with othergovernment agencies to tokenize some of the otherpersonally identifiable information such as yearlyincome, social security number etc. in the sameway, allowing you to connect the data moreprecisely with additional data.Data tokenization’s most powerful application islikely this mingling of tokenized government datawith other data sources to generate powerfulinsights—securely and with little risk to privacy(figure 1). Apart from the ability to deidentifystructured data, tokenization can even be used tode-identify and share unstructured data. Asgovernments increasingly use such data,tokenization offers many new use cases for sharingdata that resides in emails, images, text files, andother such files.</OtherInformation></Objective><Objective><Name>Financial Data</Name><Description>Keep financial data safe</Description><Identifier>_cbb6ffd0-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Data tokenization already helps some companies keep financial data safe -- Data tokenization is already considered a proventool by many. It is widely used in the financialservices industry, particularly for credit cardprocessing. One research firm estimates that thedata tokenization market will grow from US$983million in 2018 to US$2.6 billion by 2023,representing a compound annual growth rate of22 percent.1It’s not hard to understand why data tokenizationappeals to those who deal with financialinformation. Online businesses, for instance, wantto store payment card information to analyzecustomer purchasing patterns, develop marketingstrategies, and for other purposes. To meet thePayment Card Industry Data Security Standard(PCI DSS) for storing this information securely, acompany needs to put it on a system with strongdata protection. This, however, can be expensive—especially if the company must maintain multiplesystems to hold all the information it collects.Storing the data in tokenized form can allowcompanies to meet the PCI DSS requirements at alower cost compared to data encryption.2 (See thesidebar “The difference between encryption andtokenization” for a comparison of the twomethods.)Instead of saving the real card data, businessessend it to a tokenization server that replaces the actual card data with a tokenized version, savingthe key file to a secure data vault. A company canthen use the tokenized card information for avariety of purposes without needing to protect it toPCI DSS standards. All that needs this level ofprotection is the data vault containing the key file,which would be less expensive than working tosecure multiple systems housing copies of realcredit card numbers.3</OtherInformation></Objective><Objective><Name>Opiods</Name><Description>Address the opioid crisis</Description><Identifier>_cbb708d6-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Policymakers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Doctors</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Hospitals</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Insurers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Law Enforcement</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Treatment Centers</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Public Health Agencies</Name><Description>Government health data is difficult to share—as itshould be. Various agencies house large amountsof sensitive data, including both personallyidentifiable information (PII) and personal healthinformation (PHI). Given government’s significantrole in public health through programs such asMedicare, Medicaid, and the Affordable Care Act,US government agencies must expend considerableresources in adhering to Health InsurancePortability and Accountability Act (HIPAA)regulations. HIPAA alone specifies 18 different types of PHI, including social security numbers,names, addresses, mental and physical healthtreatment history, and more.5</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>US Department of Health and Human Services</Name><Description>However, the US Department of Health andHuman Services (HHS) guidelines for HIPAA notethat these restrictions do not apply to de-identifiedhealth information: “There are no restrictions onthe use or disclosure of de-identified healthinformation. De-identified health informationneither identifies nor provides a reasonable basisto identify an individual.”6</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>States</Name><Description>By tokenizing data,states may be able to share opioid-related dataoutside the agency or organization that collected itand combine it with other external data—eitherfrom other public agencies or third-party datasets—to gain missioncriticalinsights.Data tokenizationmight enablestates to bringtogether opioid-relateddata fromvarious governmentsources—including health care,child welfare, and law enforcement agencies—andcombine this data with publicly available datarelated to the social determinants of health andhealth behaviors. The goal would be to gaininsights into the causes and remedies of opioidabuse disorder.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>By tokenizing the data’s personalinformation, including PII and PHI, governmentagencies can share sensitive but critical data onopioid use and abuse without compromisingprivacy. Moreover, only the government agencythat owns the sensitive data in the first place wouldbe able to reidentify (detokenize) that data,assuming that the matching key file never leavesthat agency’s secure control. At no point in theentire cycle should any other agency or third partybe able to see real data.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Private Sector</Name><Description>The opioid crisis is not government’s onlyecosystem challenge. As we’ve seen in the privatesector, in industries from retail to auto insurance,more information generally means betterpredictions.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Retail Sector</Name><Description/></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Auto Insurers</Name><Description> (That’s why your auto insurer wants toknow about your driving patterns, and why theymight offer a discount if you ride around with anapp that captures telemetrics.)</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Companies</Name><Description>Many companies’websites require an opt-in agreement to allowthem to use an individual’s data.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Citizens</Name><Description>This approach ismore challenging in government, however, due tothe sensitive nature of data that is collected and thefact that citizens must be served whether they optin or not. Where circumstances make it impracticalto obtain consent, data tokenization can make itpossible for governments to use data in ways otherthan the originally intended use without violatingindividuals’ privacy.</Description></Stakeholder><OtherInformation>How data tokenization could help government address the opioid crisis -- Policymakers often describe the US opioid crisis asan “ecosystem” challenge because it involves somany disparate players: doctors, hospitals,insurers, law enforcement, treatment centers, andmore. As a result of this proliferation of players,information that could help tackle the problem—much of it of a sensitive nature—is held in manydifferent places...Why not simply use completely anonymized datato investigate sensitive topics like opioid use? Onereason is that tokenized data, but not anonymizeddata, can provide insights at the individual level aswell as at the aggregate level—but only to thosewho have access to the key file. For example,tokenization can turn the real Jane Jones into“Sally Smith,” allowing an agency to collectadditional data about “Sally.” If we know that“Sally Smith” is a 45-year-old female with diabetesfrom a certain zip code, the agency can merge thatwith information from hospital records about thelikelihood of middle-aged females requiringreadmission, or about the likelihood of a personfailing to follow his or her medication regimen. Ananalysis of this combined information can allowthe agency to come up with a predictive score—andthe agency can then detokenize “Sally Smith” todeliver customized insights for the very real JaneJones. This ability to gain individual-level insightscould be helpful both in delivering targetedservices and inreducingimproperlyawarded benefitsthrough fraudand abuse.The mechanicsof securelycombiningdifferent data sets after tokenization can becomplicated, but the potential benefitsare immense.</OtherInformation></Objective><Objective><Name>Other Uses</Name><Description>Take data tokenization beyond data-sharing</Description><Identifier>_cbb70e80-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.4</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>Taking data tokenization beyond data-sharing -- Beyond allowing agencies to share data withoutcompromising privacy, data tokenization can helpgovernments in other ways as well. Three potentialuses include developing and testing new software,supporting user training and demos, and securingarchived data.</OtherInformation></Objective><Objective><Name>Software</Name><Description>Developing and testing new software.</Description><Identifier>_cbb714d4-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.4.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Software Developers</Name><Description>Developers need data to build and rigorously testnew applications.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>Government frequently relies onthird-party vendors to build such systems, becauseit can be prohibitively expensive to require alldevelopment and testing be done in-house ongovernment premises. But what about systemssuch as those relating to Unemployment Insuranceor Medicaid, which contain a great deal of PII andPHI?</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Third-Party Developers</Name><Description>By using format-preserving tokens totokenize actual data while maintaining theoriginal format—where a tokenized birth date, forexample, looks like 03/11/1960 and notXY987ABC—third-party developers can work withdata that “feels” real to the system, reliablymimicking the actual data without requiring all thesecurity that would be needed if actual data wasbeing shared.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>States</Name><Description>Some US states, including Colorado,have used data tokenization in this manner. Datatokenization apps are often a cost-effective way togive developers tokenized data.</Description></Stakeholder><Stakeholder StakeholderTypeType="Organization"><Name>Colorado</Name><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Training &amp; Demos</Name><Description>Create training environments using tokenized data</Description><Identifier>_cbb71eac-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.4.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>User training and demos. When newemployees join government agencies, they oftenundergo a probationary period during which theyneed to be trained on various applications andevaluated on their performance.7 During this time,government agencies can create trainingenvironments using tokenized data, enabling newhires to work and interact with data that looks realbut does not compromise security.</OtherInformation></Objective><Objective><Name>Archived Data</Name><Description>Secure archived data. </Description><Identifier>_cbb72514-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.4.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>US Government Agencies</Name><Description>For US government agencies, securingsensitive data not in active use in a production environment has been a challenge due to costs andcompeting priorities.8</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Office of Management and Budget</Name><Description>A 2018 report by the USOffice of Management and Budget found that,while 73 percent of US agencies have prioritizedand secured their data in transit, less than16 percent of them were able to secure their data atrest9—an alarming statistic, considering thatgovernments are often a prime target forcyberattacks and have been experiencingincreasingly complex and sophisticated databreaches in the last few years.10</Description></Stakeholder><OtherInformation>Data tokenization canalso allow governments to archive sensitive dataoffsite.</OtherInformation></Objective><Objective><Name>Starting</Name><Description>Get started</Description><Identifier>_cbb72b68-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.5</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name>Public Officials</Name><Description>Public officials are often frustrated by their lack ofability to share data across organizationalboundaries, even in situations where sharing thedata would have a clear benefit. This lack of crossagencysharing can mean that agencies don’t makethe most of data analytics that could improvepublic health, limit fraud, and make betterdecisions.</Description></Stakeholder><Stakeholder StakeholderTypeType="Generic_Group"><Name>Government Agencies</Name><Description>Data tokenization can be one way forgovernment agencies to share information withoutcompromising privacy. Though it is no magic bullet,the better insights that can come from sharingtokenized data can, in many circumstances, helpgovernments achieve better outcomes.</Description></Stakeholder><OtherInformation>How can governments get started?Figure 2 depicts three important decisionsgovernment departments should carefully considerto successfully implement data tokenization. First,when should they use tokenization? Second, whatdata should they tokenize? And third, how will theysafeguard the token keys? Once the decision to usetokenization has been made, there is still muchimportant work to be done. The team tokenizingthe data must work closely with data experts toensure that tokenization is done in a way thatallows the end users to meet their intendedobjectives yet ensures privacy.</OtherInformation></Objective><Objective><Name>Use Cases</Name><Description>Decide when to use tokenization</Description><Identifier>_cbb734be-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.5.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>* Merged data analytics* Cross-agency data-sharing* Third-party development* Data archiving</OtherInformation></Objective><Objective><Name>Data</Name><Description>Decide what to tokenize</Description><Identifier>_cbb73aa4-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.5.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Data Fields</Name><Description>Identify which data fields contain sensitive information</Description><Identifier>_cbb74300-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.5.2.1</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Database Locations</Name><Description>Determine all locations where identified fields are housed in the database</Description><Identifier>_cbb74cf6-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.5.2.2</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Methodology</Name><Description>Determine how to tokenize: partial tokenization or complete tokenization</Description><Identifier>_cbb755ac-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.5.2.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation/></Objective><Objective><Name>Keys</Name><Description>Decide how to safeguard the token keys</Description><Identifier>_cbb75b7e-374d-11ea-aeb0-d62d0a83ea00</Identifier><SequenceIndicator>10.5.3</SequenceIndicator><Stakeholder StakeholderTypeType="Generic_Group"><Name/><Description/></Stakeholder><OtherInformation>* In-house on a local server* Hosted on cloud* Third-party token provider</OtherInformation></Objective></Goal></StrategicPlanCore><AdministrativeInformation><StartDate/><EndDate/><PublicationDate>2020-01-14</PublicationDate><Source>https://documents.deloitte.com/insights/CDOplaybook</Source><Submitter><GivenName>Owen</GivenName><Surname>Ambur</Surname><PhoneNumber/><EmailAddress>Owen.Ambur@verizon.net</EmailAddress></Submitter></AdministrativeInformation></PerformancePlanOrReport>
