Showing posts with label Web Programming. Show all posts
Showing posts with label Web Programming. Show all posts
Data security

Data security

Data security technologies

Disk encryption

Disk encryption refers to encryption technology that encrypts data on a hard disk drive. Disk encryption typically takes form in either software (see disk encryption software) or hardware (see disk encryption hardware). Disk encryption is often referred to as on-the-fly encryption (OTFE) or transparent encryption.

Software versus hardware-based mechanisms for protecting data

Software-based security solutions encrypt the data to protect it from theft. However, a malicious program or a hacker could corrupt the data in order to make it unrecoverable, making the system unusable. Hardware-based security solutions can prevent read and write access to data and hence offer very strong protection against tampering and unauthorized access.
Hardware based security or assisted computer security offers an alternative to software-only computer security. Security tokens such as those using PKCS#11 may be more secure due to the physical access required in order to be compromised. Access is enabled only when the token is connected and correct PIN is entered (see two-factor authentication). However, dongles can be used by anyone who can gain physical access to it. Newer technologies in hardware-based security solves this problem offering fool proof security for data.
Web Search Engine.

Web Search Engine.

A world-wide-web search results is a software program system in which was designed to seek out information on virtual reality. The particular google search are often displayed inside a distinct outcomes often referred to as yahoo and google webpages (SERPs). The knowledge may be a mix of websites, photos, and also other kinds of data files. Many search engines like google furthermore quarry facts for sale in databases or maybe open directories. In contrast to world-wide-web directories, which might be maintained just by simply individual authors, search engines like google furthermore keep real-time info by simply managing the algorithm on the world-wide-web crawler.

History.

Through early on advancement of the world wide web, there was clearly a list of webservers edited by Bob Berners-Lee and also hosted for the CERN webserver. A single fantastic snapshot of the listing in 1992 is still, nevertheless as more and more webservers proceeded to go on the web the actual middle listing can no longer continue. On the NCSA internet site, brand-new servers have been announced underneath the subject "What's Brand new! "

The primary application used by looking on the net ended up being Archie. The title symbolizes "archive" without worrying about "v". It had been created in 1990 by Joe Emtage, Expenses Heelan and also N. Andrew d Deutsch, pc science students at McGill University or college in Montreal. This system downloadable the actual directory site entries of all records on public unknown FTP (File Transport Protocol) web sites, setting up a searchable data bank associated with record names; nevertheless, Archie wouldn't directory the actual contents of the web sites because how much info ended up being therefore minimal it could be easily searched manually.

The climb associated with Gopher (created in 1991 by Tag McCahill for the University or college associated with Minnesota) generated a pair of brand-new search software programs, Veronica and also Jughead. Including Archie, they will searched the actual record names and also games kept in Gopher directory techniques. Veronica (Very Quick Rodent-Oriented Net-wide Index to be able to Computerized Archives) furnished a key word search of all Gopher menus games from the overall Gopher entries. Jughead (Jonzy's Widespread Gopher Chain of command Excavation And also Display) ended up being one tool with regard to obtaining menus info from distinct Gopher servers. Even though the title of the internet search engine "Archie" had not been a mention of the the actual Archie comic book series, "Veronica" and also "Jughead" tend to be characters from the series, so referencing their own precursor.
Within the summer time connected with 1993, not any search results existed for the world-wide-web, although many specialized inventories have been managed manually. Oscar Nierstrasz for the College connected with Geneva authored some Perl scripts that will frequently mirrored most of these websites and rewrote these people right regular formatting. This specific shaped the cornerstone pertaining to W3Catalog, the particular web's initial medieval search results, launched on September two, 1993.

With July 1993, Matthew Gray, after that in MIT, created the concepts really the initial world-wide-web trading program, the particular Perl-based Web Wanderer, and utilized this to build a great listing known as 'Wandex'. The purpose of the particular Wanderer seemed to be for you to calculate the dimensions of the internet, which in turn this would until finally later 1995. The web's second search results Aliweb shown up inside December 1993. Aliweb wouldn't work with a world-wide-web trading program, but rather counted on staying notified through web site administrators with the lifestyle in every site of an listing document in a very unique formatting.

JumpStation (created inside 12 1993 through Jonathon Fletcher) utilized some sort of world-wide-web trading program to uncover websites and develop the listing, and utilized some sort of world-wide-web type since the program for you to the problem plan. It turned out as a result the very first WWW resource-discovery application to combine the particular three essential highlights of some sort of world-wide-web search results (crawling, indexing, and searching) seeing that referred to underneath. Due to confined resources on the particular platform this went on, the indexing and therefore researching have been restricted to the particular game titles and titles seen in the web websites the particular crawler encountered.
One of many initial "all text" crawler-based search engines like google seemed to be WebCrawler, which in turn arrived on the scene inside 1994. Unlike the predecessors, this allowed people to look for any kind of concept in different website, which in turn is among the most regular for many main search engines like google due to the fact. It turned out also the very first 1 reputed with the community. Furthermore inside 1994, Lycos (which commenced in Carnegie Mellon University) premiered and grew to become an essential business project.

Soon after, quite a few search engines like google shown up and vied pertaining to acceptance. These kind of bundled Magellan, Inspire, Infoseek, Inktomi, North Gentle, and AltaVista. Google! seemed to be very favorite approaches for folks to uncover websites connected with curiosity, nevertheless the search operate operated on the world-wide-web directory site, as an alternative to the full-text copies connected with websites. Facts searchers might also browse the directory site rather than conducting a keyword-based search.

Yahoo and google followed the concept of marketing search engine terms inside 1998, from your modest search results organization referred to as goto. com. This specific transfer got a large relation to the particular SONY ERICSSON organization, which in turn journeyed through fighting for you to the most rewarding businesses inside the world-wide-web.

With 1996, Netscape seemed to be trying to provide a solitary search results a unique work since the showcased search results on Netscape's browser. There seemed to be a great deal curiosity that will rather Netscape struck handles all 5 with the main search engines like google: pertaining to $5 mil 1 year, every search results could be inside revolving within the Netscape search results page. The all 5 motors have been Google!, Magellan, Lycos, Infoseek, and Inspire.
Engines like google have been generally known as a lot of the most able minded celebrities inside the Internet making an investment frenzy that will occurred inside the later 1990s. Several businesses entered this market stunningly, having file benefits during their first community promotions. Some took straight down their particular community search results, and are also advertising and marketing enterprise-only features, including North Gentle. A lot of search results businesses have been swept up inside the dot-com bubble, some sort of speculation-driven marketplace rate of growth that will peaked inside 1999 and concluded inside 2001.

All around 2000, Google's search results went up for you to dominance. This company attained superior outcomes for most looks with an advancement known as Pr juice, seeing that seemed to be spelled out inside the paper Anatomy connected with search engines authored by Sergey Brin and Larry Site, the particular afterwards proprietors connected with Yahoo and google. This specific iterative algorithm rankings websites while using quantity and Pr juice connected with different sites and websites that will link right now there, within the idea that will great or maybe desirable websites are generally linked to a lot more than some others. Yahoo and google also managed some sort of minimalist program for you to the search results. In contrast, many of the opposition embedded search engines in a very world-wide-web site. Actually, Yahoo and google search results grew to become consequently favorite that will spoof motors come forth including Thriller Seeker.

By means of 2000, Google! seemed to be providing search companies according to Inktomi's search results. Google! bought Inktomi inside 2002, and Overture (which possessed All of the World-wide-web and Alta Vista) inside 2003. Google! traded for you to Google's search results until finally 2004, when this launched a unique search results while using blended technology connected with the acquisitions.
Microsof company initial launched MSN Seek inside late 1998 employing search engine results through Inktomi. With first 1999 the web page started to display entries through Looksmart, blended thoroughly with outcomes through Inktomi. Intended for awhile inside 1999, MSN Seek utilized outcomes through AltaVista have been rather. With 2004, Microsof company started some sort of move for you to a unique search engineering, powered through a unique world-wide-web crawler (called msnbot).

Microsoft's rebranded search results, Msn, premiered on July 1, '09. Upon September 29, '09, Google! and Microsof company selected some sort of work during which Google! Seek could be powered through Microsof company Msn engineering.

How Web Search Engines Work.

A search engine operates in the following order:
  1. Web crawling
  2. Indexing
  3. Searching
Internet search engines like yahoo function simply by stocking details about a lot of web pages, that they get from the HTML markup of the web pages. These kinds of web pages are usually reclaimed by a Internet crawler (sometimes generally known as a new spider) — the programmed Internet crawler which uses just about every website link on the webpage. The web page seller could rule out unique web pages by utilizing bots. txt.

This search results subsequently analyzes your subject matter of each and every page to view how it should be found (for case, text might be taken from the game titles, page content, headings, or even specific grounds called meta tags). Information concerning web pages are usually located in an index data source for use within in the future requests. A new query coming from a end user can be a one concept. This index allows come across data concerning your query at once. Several search engines like yahoo, like Yahoo and google, shop many or even part of the source page (referred in order to as a cache) in addition to details about online web pages, while people, like AltaVista, shop just about every concept of the page they will come across. [citation needed] This particular cached page usually holds the exact look for word considering that it's the one which ended up being really found, consequently it may be invaluable if your content of the latest page may be up-to-date plus the keyword phrases are usually will no longer inside. This issue might often be a gentle form of linkrot, in addition to Google's coping with of the usb ports improves simplicity simply by enjoyable end user anticipation how the keyword phrases will be around the went back web site. This particular fulfills your process involving very least astonishment, since end user usually needs how the keyword phrases will be around the went back web pages. Greater look for relevance can make these kind of cached web pages invaluable when they might consist of facts that could will no longer be around elsewhere.
Whenever a end user gets into a new query into yahoo search (typically by utilizing keywords), your engine investigates their index and provides all of the best-matching web pages in line with their criteria, generally using a short synopsis containing your document's title in addition to often areas of the written text. This index is created from the data located using the facts plus the way the knowledge will be found. By 2007 your Yahoo and google. com search results provides helped you to definitely look for simply by night out simply by pressing "Show look for tools" inside leftmost column of the primary listings page, and picking the desired night out range. [citation needed] The majority of search engines like yahoo assistance the employment of your boolean staff AS WELL AS, OR EVEN and never to help expand establish your look for query. Boolean staff are usually for literal queries which let the end user in order to improve in addition to expand your conditions of the look for. This engine seeks what or even words just as joined. Several search engines like yahoo offer an superior attribute called area look for, that allows consumers in order to specify the distance among search phrases. There is concept-based searching the place that the investigation entails using statistical investigation about web pages containing what or even words a person hunt for. As well, organic vocabulary requests let the end user in order to sort a new question inside similar form you should inquire the item into a human being. A site like this will be inquire. com.
This convenience involving yahoo search is determined by your relevance of the end result established the item presents returning. Although there could be an incredible number of web pages which include a certain concept or even term, some web pages might be far more related, well-known, or even well-respected than people. The majority of search engines like yahoo utilize solutions to position the final results to deliver your "best" results first. Just how yahoo search determines which web pages will be the ideal suits, in addition to exactly what purchase the final results needs to be revealed throughout, ranges widely derived from one of engine to a different. Particularly furthermore adjust after some time seeing that World-wide-web use alterations in addition to brand-new tactics develop. You'll find 2 primary sorts of search results that have progressed: you are a method involving predefined in addition to hierarchically requested search phrases which mankind have hard-wired substantially. Another is really a process which generates the "inverted index" simply by inspecting text messages the item finds. This particular first form is based much more heavily on my pc itself to try and do the bulk of the project.

The majority of Internet search engines like yahoo are usually commercial endeavors helped simply by promotion revenue therefore a variety of them allow advertisers to get the bookings positioned greater searching results for a charge. Search engines like yahoo that do not really acknowledge cash for his or her listings earn cash simply by managing look for linked advertising with the off the shelf position in search results. The major search engines earn cash each and every time somebody clicks about one of these advertising.

Market Share.

Google is the world's most popular search engine, with a marketshare of 66.44 percent as of December, 2014. Baidu comes in at second place.
The world's most popular search engines are:
Search engineMarket share in December 2014
Google66.44%
 
Baidu11.15%
 
Bing10.29%
 
Yahoo!9.31%
 
AOL0.53%
 
Ask0.21%
 
Lycos0.01%
 
Search engineMarket share in October 2014
Google58.01%
 
Baidu29.06%
 
Bing8.01%
 
Yahoo!4.01%
 
AOL0.21%
 
Ask0.10%
 
Excite0.00%

Search engineMarket share in July 2014
Google68.69%
 
Baidu17.17%
 
Yahoo!6.74%
 
Bing6.22%
 
Excite0.22%
 
Ask0.13%
 
AOL0.13%
 

Search Engine bias.

While search engines are generally designed to help get ranking web sites according to a number of combination of their own reputation and also relevancy, empirical reports reveal a variety of political, financial, and also interpersonal biases inside details they furnish. These kind of biases might be the result of financial and also industrial processes (e. grams., businesses that will market along with the search engines could become likewise most liked in its organic and natural lookup results), and also political processes (e. grams., the removal of search engine results to help conform to local laws). One example is, Yahoo will not likely surface area particular Neo-Nazi web sites in Italy and also Indonesia, where by Holocaust refusal is usually unlawful.

Biases can even be a consequence of interpersonal processes, seeing that internet search engine algorithms are frequently made to rule out non-normative viewpoints in support of additional "popular" final results. Indexing algorithms regarding major search engines skew toward insurance policy coverage regarding Ough. Utes. -based internet sites, rather than web sites through non-U. Utes. nations around the world.

Yahoo Bombing is usually one of these of your make an effort to manipulate search engine results for political, interpersonal or industrial causes.Customized Results And Filter Bubbles.

Customized Results And Filter Bubbles.

Several search engines for example The search engines in addition to Yahoo offer tailored benefits good owner's action history. This kind of causes an effect that's been known as a filtering bubble. The term talks about a sensation in which internet sites utilize algorithms in order to selectively do you know what facts a user would choose to see, based on info on the person (such since spot, previous press actions in addition to search history). Consequently, internet sites tend to demonstrate only facts which agrees with this owner's previous viewpoint, properly identifying the person inside a bubble which tends to leave out counter facts. Perfect cases usually are Google's personal search engine results in addition to Facebook's personal information stream. As outlined by Eli Pariser, whom coined the term, consumers obtain less experience of inconsistent views and are also singled out intellectually into their unique informative bubble. Pariser linked a good example in which just one user looked for The search engines regarding "BP" in addition to obtained purchase information with regards to United kingdom Oil while yet another searcher obtained info on this Deeply water Horizon oil overflow knowning that both search engine results web pages had been "strikingly different". The particular bubble consequence might have bad significance regarding civic discourse, based on Pariser.
Considering that this concern continues to be identified, competing search engines include come forth which search for in order to avoid this concern by simply not really pursuing as well as "bubbling" consumers.

Faith-Based Search Engines.

The particular world-wide increase in the Net and popularity associated with electronic material inside the Arab-speaking and Muslim Globe over the last 10 years provides inspired trust adherents, particularly in the centre Eastern and Oriental sub-continent, for you to "dream" of these individual faith-based my partner and i. at the. "Islamic" yahoo and google as well as television look for web sites filtration that could help consumers to stop opening a no-no web sites for example pornography and would likely simply enable them to gain access to web-sites that are agreeable to the Islamic trust. Quickly prior to Muslim simply 30 days associated with Ramadan, Halalgoogling that collects benefits from various other engines like google and Bing seemed to be launched to the entire world This summer 2013 for you to presents the actual halal leads to their consumers, almost 2 yrs right after I’mHalal, one more search engine optimization to begin with (launched with Sept 2011) for you to provide Midst Eastern Net was required to shut their look for assistance caused by just what their owner held responsible with not enough funding.

Even though not enough purchase and gradual velocity inside technologies inside the Muslim Globe because primary consumers as well as targeted owners provides obstructed development and thwarted success associated with serious Islamic search engine optimization, the actual amazing failure associated with seriously put in Muslim way of living web projects just like Muxlim, that gotten vast amounts from people just like Ceremony Net Endeavors, provides -- based on I’mHalal shutdown observe -- created almost laughable taking that approach how the next Facebook or myspace as well as Yahoo and google can certainly simply result from the middle Eastern in the event you assistance your current bright junior. However Muslim world wide web authorities have been figuring out for years what exactly is as well as is just not allowed based on the "Law associated with Islam" and possess recently been categorizing web sites and this sort of in to getting often "halal" as well as "haram". All of the existing and earlier Islamic yahoo and google are merely tailor made look for indexed as well as monetized by simply web major look for titans just like Yahoo and google, Yahoo and Bing along with simply selected selection methods applied in order that his or her consumers can not gain access to Haram web-sites, such as this sort of web-sites as nudity, gay, gambling as well as something that can be deemed to be anti-Islamic.

Yet another religiously-oriented search engine optimization can be Jewogle, that's the actual Jewish version associated with Yahoo and google but one more can be SeekFind. org, a Christian internet site which includes filtration blocking consumers from experiencing something on-line that will assaults as well as degrades his or her trust.
Joomla.

Joomla.

Joomla is a free and open-source content management system (CMS) for publishing web content. It is built on a model–view–controller web application framework that can be used independently of the CMS. Joomla is written in PHP, uses object-oriented programming (OOP) techniques (since version 1.5) and software design patterns,stores data in a MySQL, MS SQL (since version 2.5), or PostgreSQL (since version 3.0) database, and includes features such as page caching, RSS feeds, printable versions of pages, news flashes, blogs, polls, search, and support for language internationalization. As of February 2014, Joomla has been downloaded over 50 million times. Over 7,700 free and commercial extensions are available from the official Joomla! Extension Directory, and more are available from other sources. It is estimated to be the second most used content management system on the Internet after WordPress. 

History.

Joomla was the result of a fork of Mambo on August 17, 2005. At that time, the Mambo name was a trademark of Miro International Pvt. Ltd., who formed a non-profit foundation with the stated purpose of funding the project and protecting it from lawsuits. The Joomla development team claimed that many of the provisions of the foundation structure violated previous agreements made by the elected Mambo Steering Committee, lacked the necessary consultation with key stakeholders and included provisions that violated core open source values. Joomla developers created a website called Open Source Matters.org (OSM) to distribute information to the software community. Project leader Andrew Eddie wrote a letter that appeared on the announcements section of the public forum at mamboserver.com. Over one thousand people joined Open Source Matters.org within a day, most posting words of encouragement and support. The website received the Slashdot effect as a result. Miro CEO Peter Lamont responded publicly to the development team in an article titled "The Mambo Open Source Controversy — 20 Questions With Miro". This event created controversy within the free software community about the definition of open source. Forums of other open-source projects were active with postings about the actions of both sides. In the two weeks following Eddie's announcement, teams were re-organized and the community continued to grow. Eben Moglen and the Software Freedom Law Center (SFLC) assisted the Joomla core team beginning in August 2005, as indicated by Moglen's blog entry from that date and a related OSM announcement. The SFLC continue to provide legal guidance to the Joomla project. On August 18, Andrew Eddie called for community input to suggest a name for the project. The core team reserved the right for the final naming decision, and chose a name not suggested by the community. On September 22, the new name, Joomla!, was announced. It is the anglicised spelling of the Swahili word jumla meaning all together or as a whole which also has a similar meaning in at least Amharic, Arabic and Urdu. On September 26, the development team called for logo submissions from the community and invited the community to vote on the logo; the team announced the community's decision on September 29. On October 2, brand guidelines, a brand manual, and a set of logo resources were published. Joomla won the Packt Publishing Open Source Content Management System Award in 2006, 2007, and 2011. On October 27, 2008, PACKT Publishing announced that Johan Janssens was the Most Valued Person (MVP), for his work as one of the lead developers of the 1.5 Joomla Framework and Architecture. In 2009 Louis Landry received the Most Valued Person award for his role as Joomla architect and development coordinators 

Version History.

oomla 1.0 was released on September 22, 2005 as a rebranded release of Mambo 4.5.2.3 that combined other bug and moderate-level security fixes. Joomla 1.5 was released on January 22, 2008, and the latest release of this version was 1.5.26 on March 27, 2012.This version was the first to attain long-term support (LTS); such versions are released each three major or minor releases and supported until three months after the next LTS version is released. April 2012 marks the official end-of-life of Joomla 1.5; with Joomla 3.0 released, support for Joomla 1.5 faded away in April 2013. Joomla 1.6 was released on January 10, 2011. This version adds a full access control list functionality plus, user-defined category hierarchy, and admin interface improvements. Joomla 1.7 was released on July 19, 2011, six months after 1.6.0. This version adds enhanced security and improved migration tools. Joomla 2.5 was released on January 24, 2012, six months after 1.7.0. This version is a long term support (LTS) release. Originally this release was to be 1.8.0, however the developers announced August 9 that they would rename it to fit into a new version number scheme in which every LTS release is an X.5 release. This version was the first to run on other databases besides MySQL. Support for this version was extended until the end of 2014. Joomla 3.0 was released on September 27, 2012. Originally, it was supposed to be released in July 2012; however, the January/July release schedule was uncomfortable for volunteers, and the schedule was changed to September/March releases. On December 24, 2012, it was decided to add one more version (3.2) to the 3.x series to improve the development life cycle and extend the support of LTS versions. This will also be applied to the 4.x series. Joomla 3.1 was released on April 24, 2013. Release 3.1 includes several new features including tagging. Joomla 3.2 was released on November 6, 2013. Release 3.2 highlighting Content Versioning. Joomla 3.3 was released on April 30, 2014. Release 3.3 features improved password hashing and micro data and documentation powered by MediaWiki Translate extension. 

Deployment.

Like many other web applications, Joomla may be run on a LAMP stack.
Many web hosts have control panels for automatic installation of Joomla. On Windows, Joomla can be installed using the Microsoft Web Platform Installer, which automatically detects and installs dependencies, such as PHP or MySQL. Many web sites provide information on installing and maintaining Joomla sites. 

Extensions.

Joomla extensions extend the functionality of Joomla websites. Five types of extensions may be distinguished: components, modules, plugins, templates, and languages. Each of these extensions handles a specific function.
  • Components are the largest and most complex extensions. Most components have two parts: a site part and an administrator part. Every time a Joomla page loads, one component is called to render the main page body. Components produce the major portion of a page because a component is driven by a menu item.
  • Plugins are advanced extensions and are, in essence, event handlers. In the execution of any part of Joomla, a module or a component, an event may be triggered. When an event is triggered, plugins that are registered to handle that event execute. For example, a plugin could be used to block user-submitted articles and filter text. The line between plugins and components can sometimes be a little fuzzy. Sometimes large or advanced plugins are called components even though they don't actually render large portions of a page. An SEF URL extension might be created as a component, even though its functionality could be accomplished with just a plugin.
  • Templates describe the main design of a Joomla website. While the CMS manages the website content, templates determine the style or look and feel and layout of a site.
  • Modules render pages in Joomla. They are linked to Joomla components to display new content or images. Joomla modules look like boxes, such as the search or login module. However, they don’t require html in Joomla to work.
  • Languages are very simple extensions that can either be used as a core part or as an extension. Language and font information can also be used for PDF or PSD to Joomla conversions. 

Quick Start Joomla Package.

Quick Start Package is fully functional Joomla package which contains CMS, modules, selected templates and plugins with the configurations and the data that is used in the demo website. The sample data can also be personalized based on the template in quick start package and is different from the default Joomla! 3.x (2.5.x) package.

API For XML And History.

API For XML And History.

Simple API for XML (SAX) is a lexical, event-driven interface in which a document is read serially and its contents are reported as callbacks to various methods on a handler object of the user's design. SAX is fast and efficient to implement, but difficult to use for extracting information at random from the XML, since it tends to burden the application author with keeping track of what part of the document is being processed. It is better suited to situations in which certain types of information are always handled the same way, no matter where they occur in the document.

Pull Parsing.

Pull parsing treats the document as a series of items which are read in sequence using the Iterator design pattern. This allows for writing of recursive-descent parsers in which the structure of the code performing the parsing mirrors the structure of the XML being parsed, and intermediate parsed results can be used and accessed as local variables within the methods performing the parsing, or passed down (as method parameters) into lower-level methods, or returned (as method return values) to higher-level methods. Examples of pull parsers include StAX in the Java programming language, XML PullParser in Smalltalk, XML Reader in PHP, ElementTree.iterparse in Python, System.Xml.XmlReader in the .NET Framework, and the DOM traversal API (NodeIterator and TreeWalker). A pull parser creates an iterator that sequentially visits the various elements, attributes, and data in an XML document. Code which uses this iterator can test the current item (to tell, for example, whether it is a start or end element, or text), and inspect its attributes (local name, namespace, values of XML attributes, value of text, etc.), and can also move the iterator to the next item. The code can thus extract information from the document as it traverses it. The recursive-descent approach tends to lend itself to keeping data as typed local variables in the code doing the parsing, while SAX, for instance, typically requires a parser to manually maintain intermediate data within a stack of elements which are parent elements of the element being parsed. Pull-parsing code can be more straightforward to understand and maintain than SAX parsing code.

Document Object Model.

The Document Object Model (DOM) is an interface-oriented application programming interface that allows for navigation of the entire document as if it were a tree of node objects representing the document's contents. A DOM document can be created by a parser, or can be generated manually by users (with limitations). Data types in DOM nodes are abstract; implementations provide their own programming language-specific bindings. DOM implementations tend to be memory intensive, as they generally require the entire document to be loaded into memory and constructed as a tree of objects before access is allowed.

Data binding.

Another form of XML processing API is XML data binding, where XML data are made available as a hierarchy of custom, strongly typed classes, in contrast to the generic objects created by a Document Object Model parser. This approach simplifies code development, and in many cases allows problems to be identified at compile time rather than run-time. Example data binding systems include the Java Architecture for XML Binding (JAXB) and XML Serialization in .NET.

 

XML As Data Type.

XML has appeared as a first-class data type in other languages. The ECMAScript for XML (E4X) extension to the ECMAScript/JavaScript language explicitly defines two specific objects (XML and XMLList) for JavaScript, which support XML document nodes and XML node lists as distinct objects and use a dot-notation specifying parent-child relationships. E4X is supported by the Mozilla 2.5+ browsers (though now deprecated) and Adobe Actionscript, but has not been adopted more universally. Similar notations are used in Microsoft's LINQ implementation for Microsoft .NET 3.5 and above, and in Scala (which uses the Java VM). The open-source xmlsh application, which provides a Linux-like shell with special features for XML manipulation, similarly treats XML as a data type, using the <[ ]> notation. The Resource Description Framework defines a data type rdf:XMLLiteral to hold wrapped, canonical XML.

 

History.

XML is an application profile of SGML (ISO 8879).
The versatility of SGML for dynamic information display was understood by early digital media publishers in the late 1980s prior to the rise of the Internet. By the mid-1990s some practitioners of SGML had gained experience with the then-new World Wide Web, and believed that SGML offered solutions to some of the problems the Web was likely to face as it grew. Dan Connolly added SGML to the list of W3C's activities when he joined the staff in 1995; work began in mid-1996 when Sun Microsystems engineer Jon Bosak developed a charter and recruited collaborators. Bosak was well connected in the small community of people who had experience both in SGML and the Web. XML was compiled by a working group of eleven members, supported by a (roughly) 150-member Interest Group. Technical debate took place on the Interest Group mailing list and issues were resolved by consensus or, when that failed, majority vote of the Working Group. A record of design decisions and their rationales was compiled by Michael Sperberg-McQueen on December 4, 1997. James Clark served as Technical Lead of the Working Group, notably contributing the empty-element "<empty />" syntax and the name "XML". Other names that had been put forward for consideration included "MAGMA" (Minimal Architecture for Generalized Markup Applications), "SLIM" (Structured Language for Internet Markup) and "MGML" (Minimal Generalized Markup Language). The co-editors of the specification were originally Tim Bray and Michael Sperberg-McQueen. Halfway through the project Bray accepted a consulting engagement with Netscape, provoking vociferous protests from Microsoft. Bray was temporarily asked to resign the editorship. This led to intense dispute in the Working Group, eventually solved by the appointment of Microsoft's Jean Paoli as a third co-editor. The XML Working Group never met face-to-face; the design was accomplished using a combination of email and weekly teleconferences. The major design decisions were reached in a short burst of intense work between August and November 1996, when the first Working Draft of an XML specification was published. Further design work continued through 1997, and XML 1.0 became a W3C Recommendation on February 10, 1998.

Sources.

XML is a profile of an ISO standard SGML, and most of XML comes from SGML unchanged. From SGML comes the separation of logical and physical structures (elements and entities), the availability of grammar-based validation (DTDs), the separation of data and metadata (elements and attributes), mixed content, the separation of processing from representation (processing instructions), and the default angle-bracket syntax. Removed were the SGML declaration (XML has a fixed delimiter set and adopts Unicode as the document character set). Other sources of technology for XML were the Text Encoding Initiative (TEI), which defined a profile of SGML for use as a "transfer syntax"; and HTML, in which elements were synchronous with their resource, document character sets were separate from resource encoding, the xml:lang attribute was invented, and (like HTTP) metadata accompanied the resource rather than being needed at the declaration of a link. The Extended Reference Concrete Syntax (ERCS) project of the SPREAD (Standardization Project Regarding East Asian Documents) project of the ISO-related China/Japan/Korea Document Processing expert group was the basis of XML 1.0's naming rules; SPREAD also introduced hexadecimal numeric character references and the concept of references to make available all Unicode characters. To support ERCS, XML and HTML better, the SGML standard IS 8879 was revised in 1996 and 1998 with WebSGML Adaptations. The XML header followed that of ISO HyTime. Ideas that developed during discussion which were novel in XML included the algorithm for encoding detection and the encoding header, the processing instruction target, the xml:space attribute, and the new close delimiter for empty-element tags. The notion of well-formedness as opposed to validity (which enables parsing without a schema) was first formalized in XML, although it had been implemented successfully in the Electronic Book Technology "Dynatext" software the software from the University of Waterloo New Oxford English Dictionary Project; the RISP LISP SGML text processor at Uni-scope, Tokyo; the US Army Missile Command IADS hypertext system; Mentor Graphics Context; Interleaf and Xerox Publishing System.

 

Versions.

There are two current versions of XML. The first (XML 1.0) was initially defined in 1998. It has undergone minor revisions since then, without being given a new version number, and is currently in its fifth edition, as published on November 26, 2008. It is widely implemented and still recommended for general use. The second (XML 1.1) was initially published on February 4, 2004, the same day as XML 1.0 Third Edition, and is currently in its second edition, as published on August 16, 2006. It contains features (some contentious) that are intended to make XML easier to use in certain cases. The main changes are to enable the use of line-ending characters used on EBCDIC platforms, and the use of scripts and characters absent from Unicode 3.2. XML 1.1 is not very widely implemented and is recommended for use only by those who need its unique features. Prior to its fifth edition release, XML 1.0 differed from XML 1.1 in having stricter requirements for characters available for use in element and attribute names and unique identifiers: in the first four editions of XML 1.0 the characters were exclusively enumerated using a specific version of the Unicode standard (Unicode 2.0 to Unicode 3.2.) The fifth edition substitutes the mechanism of XML 1.1, which is more future-proof but reduces redundancy. The approach taken in the fifth edition of XML 1.0 and in all editions of XML 1.1 is that only certain characters are forbidden in names, and everything else is allowed, in order to accommodate the use of suitable name characters in future versions of Unicode. In the fifth edition, XML names may contain characters in the Balinese, Cham, or Phoenician scripts among many others which have been added to Unicode since Unicode 3.2. Almost any Unicode code point can be used in the character data and attribute values of an XML 1.0 or 1.1 document, even if the character corresponding to the code point is not defined in the current version of Unicode. In character data and attribute values, XML 1.1 allows the use of more control characters than XML 1.0, but, for "robustness", most of the control characters introduced in XML 1.1 must be expressed as numeric character references (and #x7F through #x9F, which had been allowed in XML 1.0, are in XML 1.1 even required to be expressed as numeric character references). Among the supported control characters in XML 1.1 are two line break codes that must be treated as whitespace. Whitespace characters are the only control codes that can be written directly. There has been discussion of an XML 2.0, although no organization has announced plans for work on such a project. XML-SW (SW for skunk works), written by one of the original developers of XML, contains some proposals for what an XML 2.0 might look like: elimination of DTDs from syntax, integration of namespaces, XML Base and XML Information Set (infoset) into the base standard.
The World Wide Web Consortium also has an XML Binary Characterization Working Group doing preliminary research into use cases and properties for a binary encoding of the XML infoset. The working group is not chartered to produce any official standards. Since XML is by definition text-based, ITU-T and ISO are using the name Fast Infoset for their own binary infoset to avoid confusion (see ITU-T Rec. X.891 | ISO/IEC 24824-1).

Criticism.

XML and its extensions have regularly been criticized for verbosity and complexity. Mapping the basic tree model of XML to type systems of programming languages or databases can be difficult, especially when XML is used for exchanging highly structured data between applications, which was not its primary design goal. Other criticisms attempt to refute the claim that XML is a self-describing language (though the XML specification itself makes no such claim). JSON, YAML, and S-Expressions are frequently proposed as alternatives (see Comparison of data serialization formats); which focus on representing highly structured data rather than documents, which may contain both highly structured and relatively unstructured content.