Language Asset Management (Language Asset Editor)

目次


    What is a language asset?

    There are two types of language assets in this system: “TM (translation memory)” and “Glossary”. In both cases, as a general rule, the source and translation pair (TU = Translation Unit) is counted as one line. You can register (import) your language assets in the following three ways. For details, see User Help on Language Asset List screen. The registered language assets can be used in Quick PE and LAC (Custom MT Model Management) of this system.

    (1) Import the language data (files) that you have locally.

    (2) Extract language data from available texts on the Internet (or files you have locally) by leveraging AI, and import the aligned data.

    (3) Import the stockdata provided in this system.



    Service How to use TM
    (Non Stockdata)
    TM
    (Stockdata)
    Glossary Reference  file
    Quick PE Select as a glossary     X  
    Select as a translation memory X   X  
    Use as a reference file X X X X
    Quick MT  Use as a reference file X X X X
    LAC
    (Custom MT Model Management)
    Create a custom trained model by using as corpus for the custom training X X X  
    Create a custom glossary model by setting it to a general/custom MT model     X  
    The language pairs of language assets supported by this system are English-Japanese/Japanese-English only.

    Information in Language Asset Editor screen

    On this screen, you can edit the source/target text of a language asset. At the initial view, all the TUs (rows) in the language asset are displayed.

    • String search/replace area
      • String search: Enter the string you want to search for and click [Search] button, then only those TUs in the language asset that contains the specified string in the source or target text will be displayed.
      • Detect Duplicates: Select the check box to select [Source] (matching source), [Target] (matching target), and [Both] (matching both source and target). Clicking [Search] button without specifying a string displays only those TUs that match (duplicate) the source, target, or both in the language asset. If you specify a search string, the display will be narrowed down to only those TUs that contain the string in the source or target.
      • String replacement: The entered string for search is automatically adopted as the string to be replaced (old string). If you enter the replacement string (new string) and click [Replace Text], the corresponding string in the TUs included in the search results will be replaced in bulk.
    • TU list: You can edit the source/target text separately. Unnecessary TUs can also be removed using [Delete] icon in Batch Processing menu by selecting them with the checkbox. You can also add a TU with [+] button at the right of each row.
    • Save Draft: A maximum of 60 rows can be displayed on one page. When moving to the previous or next page, use [Save Draft] button to save the edits on each page. If you don't save draft, your edits to that page will be discarded.
    • Finish: When you finish editing, click [Finish] button to complete and save your edits. When you have finished, all your edits will be reflected in the language asset.
    When a user opens Language Asset Editor, the language asset is locked and cannot be edited by another user. If the following screen appears when you open Language Asset Editor, open it in Browse Mode without unlocking it (you cannot edit the source/target text), or force the unlock. When you unlock it, you can edit it, but please note that any previously saved edits are discarded.

    Duplicate terms in the glossary

    You can register duplicate terms (bilingual data) in the glossary, but it is recommended that a term pair (source-target text) be unique, whether referenced as a glossary or incorporated into a custom glossary model. However, there are cases where you want to change the translation of each term and the translation of a compound word, as shown below.

    English 
    Japanese 
    asset 資産
    management 管理
    asset management アセットマネジメント

    If you create a custom glossary model with a glossary which includes such bilingual data, when you machine translate the source text (in this case, English) using the model, the term with the longer source length will be prioritized and the translation will be adopted. In the above example, if “asset management” appears in the source text, “アセットマネジメント” will be prioritized as the translation, so this source text will not be output as “資産管理”.

    Bilingual data in the glossary can be edited in Language Asset Editor screen after importing. It is also possible to detect duplicates in that screen. Even if there is duplication in the glossary, there is no problem for subsequent processing.
    This system assumes that glossary includes terms that do not change much in context, such as technical terms, product names, and proper nouns. For example, when creating an English-Japanese glossary, you can unify the translations by registering two source sentences (terms), singular and plural, as separate records. However, in the case of a Japanese-English glossary that is in the opposite direction, even if you register two translation records for one source, it will not be automatically used properly in machine translation. It can be said that terms whose singular/plural form changes depending on the context should not be registered in the glossary.