Welcome!

Industrial IoT Authors: Pat Romanski, Srinivasan Sundara Rajan, Liz McMillan, Elizabeth White, Scott Allen

Related Topics: Industrial IoT

Industrial IoT: Article

What Is XLIFF and Why Should I Use It?

A brief overview of the XML Localization Interchange File Format (XLIFF)

Much time, energy, and commitment are required to develop and sell successful software products and Web-based services. Most products of this type are initially developed for a specific language and locale (e.g., U.S. English). To maximize return on investment, products can be customized so they may be available to the largest possible market - the global market. This customization process is known as localization.

Localization includes not only translation of the displayed text, but also adaptation of a product to comply with a country's cultural and legal practices. Examples of cultural conventions include date/time formats, postal address formats, font sizes, appropriateness of colors, numeric or currency formats and symbols, culturally appropriate icons or graphics, etc. The diversity of software platforms and technologies means that tools and technologies that support localization are also diverse and are frequently incompatible with each other. Industry standards drive process and technology efficiencies, and OASIS XLIFF (XML Localization Interchange File Format) has emerged as a standard interchange file format for localization-related data and metadata. This article will introduce the process of localization and summarize the challenges and issues facing those who localize. It will illustrate how XLIFF addresses many of the challenges and issues with descriptions of its architecture, provide examples of how to use it in real life, and discuss how it was developed and where it goes from here.

Localization: A Brief Overview
Having a product available in a different language increases the potential market into which it can be sold. Microsoft, IBM, and Oracle earn more than 60 percent of their sales from international markets. IBM recently reported that 70 percent of Web users speak a primary language other English. Within the United States, 18 percent of the population speak a language other than English at home. The Canadian Translation Bureau has reported that Canada represents 4 to 8 percent of the global translation services market with only 0.5 percent of the world's population.

The software localization industry was born in the mid 1980s as a result of the personal computer revolution. Early projects were very engineering intensive. Software developers would usually throw the final version of the application over the proverbial wall and expect the localization team to do the rest. The user interface and functional code were in the same file, which meant that functionality problems would often be introduced during the translation process. Before translation could begin, a software engineer would first need to change the character set code pages to ensure that software could be translated into the target language. Separating code from the UI required significant software development and localization resources. Physically and logically partitioning the UI data into separate resource containers reduced the severity and frequency of functionality defects introduced by the localization process.

The introduction of the Unicode character set standard and the growing emphasis on internationalization as a standard component of the software engineering practice was a very significant benefit to reducing the time and cost of localization. Internationalization is the process of developing software in such a way that it can run in different international environments without adapting or recompiling the code. Unicode introduced a common character set code page, and it is one of the corner-stones of software internationalization. The World Wide Web Consortium's (W3C) specifications and guidelines have addressed localization issues such as the separation of code and UI, character set issues, and date and time formatting. Cultural and economic factors are motivating many organizations to design and adapt their products for delivery to the global markets.

Today global markets are not content to wait months or even weeks for locally adapted software to be made available to them, so localization must now be done in parallel with the core product development. "Throwing software over the wall" to be localized after a product is sold is no longer an option.

The Internet introduces new challenges into the localization process by enabling more complex distribution models. Products and content are rolled out simultaneously throughout the global markets. The proliferation of diverse architectural frameworks and technologies upon which Internet and e-business applications are built has resulted in different degrees of complexity in both the core development and in localization.

Throughout its history, the localization industry addressed these business challenges through improvements in tools and processes delivered by competing vendors. Their solutions addressed specific technical or process requirements, but were rarely interoperable, which often resulted in vendor lock-in. With XLIFF, the localization industry got together to find a solution to some of the challenges that the industry is now facing.

Localization Challenges Addressed by XLIFF
XLIFF has already provided solutions to difficult localization challenges, including the following.

Challenge: Many different file formats to localize - The typical localization project is composed of data stored within many unique resource formats. Building tools to support the localization of these formats is costly and inefficient.

Solution: Transform or extract all localizable resources, regardless of native format, into XLIFF containers. For example, before XLIFF existed, one major database and enterprise software vendor localized 32 unique resource formats. Some of the formats were proprietary, others were legacy, but most were industry standards (i.e., Java List Resource and Property Resource Bundles, Windows RC Data, .html, .jsp, and various XML). Additionally, each new release introduced additional resource formats, usually XML-based, but each required extensive retooling of the localization tools, which lead to quality problems and scheduling delays. To reduce the retooling work and its consequences, the core development teams were given the responsibility of extracting or transforming the new resource formats to XLIFF before handing off for translation. Three years and three major releases later, the number of unique resource formats supported by the localization tools was reduced from 32 to a much more manageable 17. Additionally, by adopting XLIFF, the tools development team was able to shift resources that had been dedicated to retooling onto building new cost-saving enhancements such as XLIFF's built-in suggested translations features. Listing 1 is an example of a Java properties file and its XLIFF representation.

Challenge: Lack of version management metadata in native resource containers - Localization projects often run concurrently with core development projects, which means that resource files will be made of up of multiple versions or milestones. Version management at the segment level (a segment is the smallest discrete unit of translatable text) is a very useful feature if your goal is to maximize the reuse of previous translations. Web content is often dynamic, and multilingual content must be kept in sync in order to maintain quality. Few if any native resource containers provide a means of tracking the version of the content.

Solution: XLIFF provides metadata structures for tracking versions of source and translated content.

Challenge: Lack of workflow metadata in native resources - During the localization process, data passes through many different hands. Data to be localized is typically externalized by software publishers and handed off to a localization service provider, who in turn may hand it over to translation subcontractors. At each stage of the process, the types of data required are unique to the particular phase (source/target text, Translation Memory, Machine Translation, Termbase, etc.). Introducing automation into a development process saves money and resources and improves quality by ensuring the reproducibility of the process. Native resource files don't usually contain mechanisms for tracking the stage of the process at which changes were introduced to the localization process.

More Stories By Peter Reynolds

Peter Reynolds is manager of the software development team at the Dublin, Ireland office of Bowne Global Solutions (BGS), the leading provider of localization and translation solution. Peter and his team are responsible for developing some of the software that BGS use to run their business, including Elcano, the online translation service, and myInfoShare, the online project collaboration workspace. Peter has been working on XLIFF since its inception and is a founder member and secretary of the OASIS XLIFF Technical Committee. He also chairs the OASIS Translation Web Services Technical Committee.

More Stories By Tony Jewtushenko

Tony Jewtushenko is the founding and present chair of the XLIFF TC as well as the director of R&D for Product Innovator (www.productinnovator.com), a Dublin, Ireland-based consultancy that provides product management, process improvement, localization, and internationalization services to software companies.
During his 23-year career, Tony was a key contributor to dozens of releases of successful commercial software products, including Lotus 1-2-3 and Notes, Oracle JDeveloper, and iDS. His multinational work and life experience spans USA, Europe, Middle East, and Asia.

Comments (0)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


@ThingsExpo Stories
Data is an unusual currency; it is not restricted by the same transactional limitations as money or people. In fact, the more that you leverage your data across multiple business use cases, the more valuable it becomes to the organization. And the same can be said about the organization’s analytics. In his session at 19th Cloud Expo, Bill Schmarzo, CTO for the Big Data Practice at EMC, will introduce a methodology for capturing, enriching and sharing data (and analytics) across the organizati...
SYS-CON Events announced today that Bsquare has been named “Silver Sponsor” of SYS-CON's @ThingsExpo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. For more than two decades, Bsquare has helped its customers extract business value from a broad array of physical assets by making them intelligent, connecting them, and using the data they generate to optimize business processes.
SYS-CON Events has announced today that Roger Strukhoff has been named conference chair of Cloud Expo and @ThingsExpo 2016 Silicon Valley. The 19th Cloud Expo and 6th @ThingsExpo will take place on November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. "The Internet of Things brings trillions of dollars of opportunity to developers and enterprise IT, no matter how you measure it," stated Roger Strukhoff. "More importantly, it leverages the power of devices and the Interne...
19th Cloud Expo, taking place November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, will feature technical sessions from a rock star conference faculty and the leading industry players in the world. Cloud computing is now being embraced by a majority of enterprises of all sizes. Yesterday's debate about public vs. private has transformed into the reality of hybrid cloud: a recent survey shows that 74% of enterprises have a hybrid cloud strategy. Meanwhile, 94% of enterpri...
In this strange new world where more and more power is drawn from business technology, companies are effectively straddling two paths on the road to innovation and transformation into digital enterprises. The first path is the heritage trail – with “legacy” technology forming the background. Here, extant technologies are transformed by core IT teams to provide more API-driven approaches. Legacy systems can restrict companies that are transitioning into digital enterprises. To truly become a lea...
According to Forrester Research, every business will become either a digital predator or digital prey by 2020. To avoid demise, organizations must rapidly create new sources of value in their end-to-end customer experiences. True digital predators also must break down information and process silos and extend digital transformation initiatives to empower employees with the digital resources needed to win, serve, and retain customers.
Businesses are struggling to manage the information flow and interactions between all of these new devices and things jumping on their network, and the apps and IT systems they control. The data businesses gather is only helpful if they can do something with it. In his session at @ThingsExpo, Chris Witeck, Principal Technology Strategist at Citrix, will discuss how different the impact of IoT will be for large businesses, expanding how IoT will allow large organizations to make their legacy ap...
Video experiences should be unique and exciting! But that doesn’t mean you need to patch all the pieces yourself. Users demand rich and engaging experiences and new ways to connect with you. But creating robust video applications at scale can be complicated, time-consuming and expensive. In his session at @ThingsExpo, Zohar Babin, Vice President of Platform, Ecosystem and Community at Kaltura, will discuss how VPaaS enables you to move fast, creating scalable video experiences that reach your...
In his keynote at 18th Cloud Expo, Andrew Keys, Co-Founder of ConsenSys Enterprise, provided an overview of the evolution of the Internet and the Database and the future of their combination – the Blockchain. Andrew Keys is Co-Founder of ConsenSys Enterprise. He comes to ConsenSys Enterprise with capital markets, technology and entrepreneurial experience. Previously, he worked for UBS investment bank in equities analysis. Later, he was responsible for the creation and distribution of life sett...
Cloud computing is being adopted in one form or another by 94% of enterprises today. Tens of billions of new devices are being connected to The Internet of Things. And Big Data is driving this bus. An exponential increase is expected in the amount of information being processed, managed, analyzed, and acted upon by enterprise IT. This amazing is not part of some distant future - it is happening today. One report shows a 650% increase in enterprise data by 2020. Other estimates are even higher....
SYS-CON Events announced today that SoftLayer, an IBM Company, has been named “Gold Sponsor” of SYS-CON's 18th Cloud Expo, which will take place on June 7-9, 2016, at the Javits Center in New York, New York. SoftLayer, an IBM Company, provides cloud infrastructure as a service from a growing number of data centers and network points of presence around the world. SoftLayer’s customers range from Web startups to global enterprises.
Internet of @ThingsExpo, taking place November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 19th Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devices - comp...
One of biggest questions about Big Data is “How do we harness all that information for business use quickly and effectively?” Geographic Information Systems (GIS) or spatial technology is about more than making maps, but adding critical context and meaning to data of all types, coming from all different channels – even sensors. In his session at @ThingsExpo, William (Bill) Meehan, director of utility solutions for Esri, will take a closer look at the current state of spatial technology and ar...
The vision of a connected smart home is becoming reality with the application of integrated wireless technologies in devices and appliances. The use of standardized and TCP/IP networked wireless technologies in line-powered and battery operated sensors and controls has led to the adoption of radios in the 2.4GHz band, including Wi-Fi, BT/BLE and 802.15.4 applied ZigBee and Thread. This is driving the need for robust wireless coexistence for multiple radios to ensure throughput performance and th...
Fifty billion connected devices and still no winning protocols standards. HTTP, WebSockets, MQTT, and CoAP seem to be leading in the IoT protocol race at the moment but many more protocols are getting introduced on a regular basis. Each protocol has its pros and cons depending on the nature of the communications. Does there really need to be only one protocol to rule them all? Of course not. In his session at @ThingsExpo, Chris Matthieu, co-founder and CTO of Octoblu, walk you through how Oct...
What are the new priorities for the connected business? First: businesses need to think differently about the types of connections they will need to make – these span well beyond the traditional app to app into more modern forms of integration including SaaS integrations, mobile integrations, APIs, device integration and Big Data integration. It’s important these are unified together vs. doing them all piecemeal. Second, these types of connections need to be simple to design, adapt and configure...
“We're a global managed hosting provider. Our core customer set is a U.S.-based customer that is looking to go global,” explained Adam Rogers, Managing Director at ANEXIA, in this SYS-CON.tv interview at 18th Cloud Expo, held June 7-9, 2016, at the Javits Center in New York City, NY.
Is your aging software platform suffering from technical debt while the market changes and demands new solutions at a faster clip? It’s a bold move, but you might consider walking away from your core platform and starting fresh. ReadyTalk did exactly that. In his General Session at 19th Cloud Expo, Michael Chambliss, Head of Engineering at ReadyTalk, will discuss why and how ReadyTalk diverted from healthy revenue and over a decade of audio conferencing product development to start an innovati...
In his general session at 18th Cloud Expo, Lee Atchison, Principal Cloud Architect and Advocate at New Relic, discussed cloud as a ‘better data center’ and how it adds new capacity (faster) and improves application availability (redundancy). The cloud is a ‘Dynamic Tool for Dynamic Apps’ and resource allocation is an integral part of your application architecture, so use only the resources you need and allocate /de-allocate resources on the fly.
What happens when the different parts of a vehicle become smarter than the vehicle itself? As we move toward the era of smart everything, hundreds of entities in a vehicle that communicate with each other, the vehicle and external systems create a need for identity orchestration so that all entities work as a conglomerate. Much like an orchestra without a conductor, without the ability to secure, control, and connect the link between a vehicle’s head unit, devices, and systems and to manage the ...