How to erase a substring that starts with a given pattern?

ERRATUM: c consists of character vectors, not strings.
c = {'Glucose C6H12O6'; 'Benzol C6H6'}
Hello,
Assume we have a cell array with string variables:
c = {"Glucose C6H12O6"; "Benzol C6H6"}
The goal is to remove chemical formulas from the strings and to get
c = {"Glucose"; "Benzol"}
The names can also be longer than one word. My idea for now is to define a pattern
pat = " C" + digitsPattern + "H" + digitsPattern;
and then to remove everything starting from that pattern till the end. I try to use the function eraseBetween, which can delete a substring between either two patterns or two positions. But my start point is a pattern, and end point - a position. So my questions are:
  1. Is there a way to flag the end of the string as a pattern that will be unique for the given string?
  2. How to find the starting position of a pattern in a string? We assume that there is only one pattern in each string.
A workaround is to add a unique pattern to the end of the string, and then delete everything between the two patterns (including them), but it is not ideal since I cannot be 100% sure the unique pattern is really unique.
Thank you in advance.

4 Comments

If you had a propoer string array then this would be much easier:
c = ["Glucose C6H12O6"; "Benzol C6H6"]
c = 2×1 string array
"Glucose C6H12O6" "Benzol C6H6"
d = regexprep(c,'\s*(\w\d+)+','')
d = 2×1 string array
"Glucose" "Benzol"
or even just get the data that you want:
d = regexp(c,'^\w+','once','match')
d = 2×1 string array
"Glucose" "Benzol"
See "Why Do Strings in Cell Arrays Return an Error?" here:
Now I am a bit embarrassed, because in fact I have a cell array of character vectors, not string ones. @Stephen23, I do not understand the syntax of your solutions yet, but I tested them, and they are not doing what I want:
c = {'Trans 4 Hydroxy L proline C5H9NO3'};
d = regexprep(c,'\s*(\w\d+)+','')
d = 1×1 cell array
{'Trans 4 Hydroxy L prolineN'}
e = regexp(c,'^\w+','once','match')
e = 1×1 cell array
{'Trans'}
The thing is, the formulas can contain other elements as well, and in some cases, the elements will not be capitalised: 'C49H55FeN4O6'.
Simpler than patterns and STRIP and all that jazz:
c = {'Trans 4 Hydroxy L proline C5H9NO3'};
d = regexprep(c,'\s*(\w+\d+)+','')
d = 1×1 cell array
{'Trans 4 Hydroxy L proline'}
Yeah, assuming one can figure out the regex pattern... :>)

Sign in to comment.

 Accepted Answer

c = ["Glucose C6H12O6"; "Benzol C6H6"];
pat = " C" + digitsPattern + "H" + digitsPattern;
extractBefore(c,pat)
ans = 2×1 string array
"Glucose" "Benzol"

3 Comments

The answer does not work for the initial, wrongly formulated, input (cell array of strings), but it works for the data that I actually have (cell array of character vectors). Therefore, I accept it; thank you!
While it is less convenient to enclose strings in a cell array, it can be dealt with if that were to actually be the case (not that I'd recomend other than switching storage)...
c = {"Glucose C6H12O6"; "Benzol C6H6"};
pat=' C'+digitsPattern+'H'+digitsPattern;
extractBefore([c{:}].',pat)
ans = 2×1 string array
"Glucose" "Benzol"
or
extractBefore(cellstr(c),pat)
ans = 2×1 cell array
{'Glucose'} {'Benzol' }
although NOTA BENE the differening result types.
As to the issue of possible case; there's a solution for that, too...
c = {"Glucose C6h12O6"; "Benzol c6H6"};
pat=caseInsensitivePattern(' C'+digitsPattern+'H'+digitsPattern);
extractBefore([c{:}],pat)
ans = 1×2 string array
"Glucose" "Benzol"
I don't know if the format of the input strings is controlled sufficiently well or not, but one other slight enhancement might be to wrap the result in a call to strip() if there were a chance of extra white space between the name and the formula--
c = cellstr(["Glucose C6h12O6"; "Benzol c6H6"]); % revert to cellstr
pat=caseInsensitivePattern(' C'+digitsPattern+'H'+digitsPattern);
extractBefore(c,pat)
ans = 2×1 cell array
{'Glucose '} {'Benzol' }
or
strip(extractBefore(c,pat))
ans = 2×1 cell array
{'Glucose'} {'Benzol' }
to avoid that extra blank at the end of the first chemical...
@dpb, thank you for the comment. Indeed, strip() is something I need for tidying the data up.

Sign in to comment.

More Answers (0)

Categories

Products

Release

R2025b

Asked:

on 22 Sep 2026 at 9:28

Commented:

dpb
on 23 Sep 2026 at 17:02

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!