Skip to content

Ambiguous IUPAC codes - #5

Open
Colelyman wants to merge 3 commits into
jts:masterfrom
Colelyman:ambiguous
Open

Colelyman wants to merge 3 commits into
jts:masterfrom
Colelyman:ambiguous

Conversation

@Colelyman

Copy link
Copy Markdown

These changes adds functionality to accept ambiguous IUPAC codes in bwtdisk_prepare.

It handles the ambiguity by choosing a random base that is within the set of bases for that ambiguity code. For example, N can be A, C, G, or T; S can be C or G; H can be A, C, or T; etc.

This required libdbgfm to be linked to bwtdisk_prepare because it uses the IUPAC methods found in alphabet.cpp.

@sjackman

sjackman commented Aug 3, 2017

Copy link
Copy Markdown
Collaborator

I'd suggest picking the lexicographically smallest possible nucleotide for that ambiguity code rather than a random one, to make the result deterministic.

@Colelyman

Copy link
Copy Markdown
Author

I have updated the function so that the lexicographically smallest possible nucleotide for each ambiguity code. Thanks for the suggestion @sjackman

Comment thread bwtdisk_prepare.cpp Outdated
}
// get the lexicographically smallest base for the code
char base = IUPAC::getPossibleSymbols(line[pos])[0];
line[pos] = base;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can omit the base intermediate variable.

line[pos] = IUPAC::getPossibleSymbols(line[pos])[0];

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed.

@sjackman

sjackman commented Aug 7, 2017

Copy link
Copy Markdown
Collaborator

What's your use case, Cole? Is it that you have reads with Ns in them, or do you have reads with other IUPAC codes in them, or are you working with sequences other than reads?

@sjackman
sjackman requested a review from jts August 7, 2017 19:37
@Colelyman

Copy link
Copy Markdown
Author

My use case is using assembled genomes. Ideally, I would like to be able to keep Ns, but then I thought it would be helpful to accept all IUPAC codes.

Do you know how hard it would be to accept Ns? I figured it might be difficult to add another character to the alphabet due to the encoding/compression.

@sjackman

sjackman commented Aug 7, 2017

Copy link
Copy Markdown
Collaborator

Jared (@jts) is in a better position to answer that question than myself.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants