<?xml version="1.0" encoding="UTF-8"?>
<PerformancePlanOrReport xmlns="urn:ISO:std:iso:17469:tech:xsd:PerformancePlanOrReport" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
 xsi:schemaLocation="urn:ISO:std:iso:17469:tech:xsd:PerformancePlanOrReport http://stratml.us/references/PerformancePlanOrReport20160216.xsd" Type="Performance_Plan"><Name>Rorur Roadmap</Name><Description>According to our current vision, there must be several stages in the development and deployment of the Distributed Search Engine (DSE). The functioning of DSE is based on splitting the workload between the participating agents, in particular, on allocation of pages to nodes. ^^ Therefore, it is necessary to first sample the web and collect a sample of urls to seed the crawl for each of the nodes. Furthermore, we must ensure that there are enough nodes to bootstrap the distribution algorithm that assigns the urls to nodes. Therefore, the first stage in the deployment will consist in autonomous operation of nodes that independently sample the web. The nodes only collect links and construct the web graph at this stage. The condition to transit to the next stage will be decided centrally by a ”monitor” node, and amounts to sampling around 10% of urls (about 109 urls).</Description><OtherInformation>We have already written the code that pertains Stage 1 and most of the code that pertains Stage 2. Now it is the issue of having enough hardware to run this project and harvest the results. Most of the software development now pertains researching and implementing ranking algorithms. In essence, this is the most interesting part of the project, when most of the dirty work is already done.</OtherInformation><StrategicPlanCore><Organization><Name>Rorur</Name><Acronym>RRR</Acronym><Identifier>_296409e4-62ab-11ed-a5a1-28b91083ea00</Identifier><Description/><Stakeholder StakeholderTypeType="Person"><Name>Stanislav Srednyak, Ph.D.</Name><Description/><Role><Name/><Description/></Role></Stakeholder></Organization><Vision><Description>A distributed search engine</Description><Identifier>_2964122c-62ab-11ed-a5a1-28b91083ea00</Identifier></Vision><Mission><Description>To specify application programming interfaces to support distributed query service nodes</Description><Identifier>_296412f4-62ab-11ed-a5a1-28b91083ea00</Identifier></Mission><Value><Name>Autonomy</Name><Description/></Value><Value><Name>Peerism</Name><Description/></Value><Goal><Name>API</Name><Description>Communicate between the monitor and the nodes</Description><Identifier>_2964138a-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>Stage 1</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>At stage 1 communication goes only between the monitor and the nodes, there is no communication between nodes. Since there are about 1010 urls currently in use, the storage of the database of these urls amount to about 100 ∗ 1010 = 1T B ( assuming about 100B per url). Note that we do not construct the web graph at this stage. Construction and maintainance of the web graph will require about 100 times more space ( estimating that there are about 100 outgoing links per page). ^^ It is very important that enough nodes join at this stage. Although some crawl can be done with small number of nodes, for example, by restricting to the most popular pages, this kind of machine would be of limited utility. ^^ We should note that with 104 nodes that operate with standard bandwidth it is possible to do a crawl of the whole web in less than a day.</OtherInformation><Objective><Name>URLs</Name><Description>Maintain local data base of the visited urls</Description><Identifier>_29641420-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api1.1</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>fetch urls, parse them, and maintain local data base of the visited urls. No url is visited twice.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba2c9a-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_1</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Data</Name><Description>Collect the url data from all nodes</Description><Identifier>_296414ac-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api1.2</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>the collected links are packed and sent to the monitor. Monitor collects the url data from all nodes.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba2fd8-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_2</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>History</Name><Description>Maintain the history of the crawl for each of the nodes</Description><Identifier>_29641542-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api1.3</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>Monitor keeps history of the crawl for each of the nodes. Once the necessary percentage is crawled, the monitor issues a signal to transit to the next stage. Node code is written so that it auto-renews itself on the command form the monitor, by fetching a new copy of the code from a pre-defined repository.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba321c-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_3</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Transfer</Name><Description>Support secure data transfer between the monitor and the nodes, and back</Description><Identifier>_296415d8-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api1.4</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>there is corresponding protocol that allows for secure data transfer between the monitor and the nodes, and back. At stage 1 the set of commands is limited to the transfer of url lists.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba3320-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_4</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective></Goal><Goal><Name>Allocation</Name><Description>Allocate pages among the nodes for analysis and ranking</Description><Identifier>_2964166e-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>Stage 2</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>This stage constitutes the core of the web analysis and ranking. At this stage the nodes receive the allocation of pages that they will be responsible for. The pages are distributed in pseudo-random fashion. No single host will be assigned to individual node. The url assignment must respect the disk space limits of individual nodes, i.e. , the url list that each node will be serving must have total size less than the disk limit specified by the user on the installation of node code. ^^ At this stage, nodes need to communicate to each other in order to ^^ 1. efficiently crawl the web ^ 2. compute the rank. ^^ At this stage, nodes will not crawl the web randomly, as at stage 1, rather, they will only crawl the assigned pages. When they encounter a link that does not belong to their name space, they will send it out to the peers that are responsible for this url. This dictates the following functionality:</OtherInformation><Objective><Name>Mapping</Name><Description>Maintain the globally available table that maps urls to nodes.</Description><Identifier>_29641768-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.1</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>On the surface of it, this looks like a large data base that must be maintained in distributed fashion.  However, there is a workaround based on hashing the urls and assigning the segments to the nodes. This results in a much smaller database, which is independent of web size, and linear in the number of nodes. We can, for example, maintain it on a central server, or even let each node keep its own copy. This last option may be the most efficient ,since it will not require fetching the node address when sending the url to the corresponding node.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba34ec-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_5</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Listening</Name><Description>Open ports to listen for incoming data</Description><Identifier>_29641812-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.2</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>Nodes must open ports to listen for incoming data that contains the urls harvested by the peers that belongs to their name segment.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba35be-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_6</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Service, Ranking &amp; Analyses</Name><Description>Serve, rank and analyze data in html pages</Description><Identifier>_296418bc-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.3</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>Nodes must serve html pages that contain the statistics of their crawl, ranking and data analysis, as well as history of the receive and send transactions ( for debugging purposes).</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba3690-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_7</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Indices</Name><Description>Construct local indices of pages</Description><Identifier>_29641966-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.4</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>Nodes construct the local index of pages in their name segment. In contrast to Stage 1, at this stage nodes save the html of the pages.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba3762-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_8</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Latency &amp; Loading</Name><Description>Collect and share latency information and load information</Description><Identifier>_29641a10-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.5</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>Nodes collect the latency information and load information for each pair, and send this information to pre-assigned monitor nodes. The monitor nodes collect this information, and based on this, calculate the roles the nodes will play in collecting information when user queries will be served. The monitor nodes are maintained by the network operators and are trusted nodes at this stage. The monitors are the trusted nodes that perform network optimization and role allocation. It is their responsibility to design the distributed index tree and return tree ( as described in the white paper) and to assign corresponding privileges to the nodes. The privilege table will be globally available. The nodes will be looking it up when verifying the url and rank information that they receive from peers.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba383e-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_9</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Ranking</Name><Description>Compute the ranking structure for pages</Description><Identifier>_29641aba-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.6</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>Nodes compute the ranking structure for the pages in their name segment. At this stage, we may limit this computation to the most primitive pagerank ( which is just a number), but the code must be written in such a way as to easily incorporate new ranking algorithms , on the request from the maintainer nodes. The defining feature of the network ranks of the pages lies in its dependence on the web graph. ^^ Therefore, this computation must be performed globally by the network, in several iterative stages, each of them consisting in evaluation of the rank based on the data received from the peers according to the incoming links for a page. It must be ensured that the rank structure is designed in such a way that this iterative computation converges in few steps ( pagerank for example has this property). We may add the feature that nodes check the existence of the incoming links , to reduce the probability of fraud. ^^ Consistency of the rank must be verified at this stage. It seems sufficient to cross check the rank computed by different nodes that cover the same name segment. However, we may add special set of pseudo-privileges for nodes to query the peers for the rank, and then check consistency at randomly chosen url. Therefore, we may add a set of functions to the node communication protocol that will allow such requests.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba3924-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_10</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective><Objective><Name>Transmission</Name><Description>Transmit information serving queries</Description><Identifier>_29641b64-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.7</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>Nodes receive part of the code that is responsible for collection and transmission of information that serves particular query. They are assigned roles in this process and receive corresponding privileges.  The tree information is globally available. In particular, nodes know the addresses for the parent nodes to which to send the merged rank lists , and the addresses and ids of the child nodes, from whom they can receive rank lists (more details see in the white paper). ^^ At this stage we will have fully functioning search engine.</OtherInformation><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba3a0a-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_11</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective></Goal><Goal><Name>Advertising</Name><Description>Implement advertising</Description><Identifier>_29641d3a-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>Stage 3</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation>At this stage, we will implement advertising system. We will detail the api later on.</OtherInformation><Objective><Name>Details</Name><Description>To be determined</Description><Identifier>_29641e0c-62ab-11ed-a5a1-28b91083ea00</Identifier><SequenceIndicator>api2.8</SequenceIndicator><Stakeholder><Name/><Description/><Role><Name/><Description/></Role></Stakeholder><OtherInformation/><PerformanceIndicator ><SequenceIndicator/><MeasurementDimension/><UnitOfMeasurement/><Identifier>_12ba3bb8-62ac-11ed-90e5-55e61083ea00</Identifier><Relationship><Identifier>PLACEHOLDER_12</Identifier><ReferentIdentifier/><Name/><Description/></Relationship><MeasurementInstance><TargetResult><Description/><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></TargetResult><ActualResult><Description>[To be determined]</Description><Descriptor><DescriptorName/><DescriptorValue/></Descriptor><StartDate/><EndDate/></ActualResult></MeasurementInstance><OtherInformation/></PerformanceIndicator></Objective></Goal></StrategicPlanCore><AdministrativeInformation><Identifier>_12ba3cc6-62ac-11ed-90e5-55e61083ea00</Identifier><StartDate>2021-07-15</StartDate><EndDate/><PublicationDate>2022-11-12</PublicationDate><Source>https://rorur.com/roadmap3.pdf</Source></AdministrativeInformation><Submitter><Identifier>_12ba3e4c-62ac-11ed-90e5-55e61083ea00</Identifier><GivenName>Owen</GivenName><Surname>Ambur</Surname><PhoneNumber/><EmailAddress>Owen.Ambur@verizon.net</EmailAddress></Submitter></PerformancePlanOrReport>