Skip to content
 
 

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OptCodeTrans

Overview

OptCodeTrans

Dataset Access

We have open-sourced both our Cangjie(仓颉) monolingual dataset and the Cangjie-Java parallel corpus. These datasets can be found in the data directory. All processed data will be released in the near future.

Training

Model training is divided into two phases: continuous pretraining and instruction fine-tuning.

  • The llm-tuning folder contains the code for continuous pretraining and instruction fine-tuning of the StarCoder model.
  • The t5-finetuning folder includes the code for instruction fine-tuning of the Code-T5p model.
  • The t5-pretraining folder contains the code for continuous pretraining of the Code-T5p model.

Evaluation

We evaluate translation results using the BLEU automated metric and Function Equivalence.

How to run?

  1. Open the bash-test/call_bash.ipynb file;
  2. Change the model_id to the model translation result you want to evaluate. For example, you can use the demo result starcoder2-3b_cangjie_it_2200_lr1e-05_ebs32 provided for easy testing.
  3. Run the notebook for the specified model translation result;
  4. In the bash-test/call_bash.ipynb notebook, you can see the evaluation results of the given demo result.

Notice

  1. the model translation result should be a jsonl file with the following format:
    {"src": "...", "pred": "..."}
    {"src": "...", "pred": "..."}
    {"src": "...", "pred": "..."}
  2. We also privide a detailed handbook on how to execute the test-based evaluation for Cangjie in English and Chinese on using the evaluation script in the bash-test folder.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages