Convert character, numeric, and date

as.character() / as.numeric() / as.Date()

SAS: PUT, INPUT, informats and formats (BEST., Zw., YYMMDD10., E8601DT.)

PUT writes a value with a format and always returns character. INPUT reads a character value with an informat. The informat sets the result type. R uses one function per target type.

Quick reference

library(dplyr)

as.character(num)                        # strip(put(num, best.))
as.numeric(chr)                          # input(chr, best.)
as.numeric(na_if(trimws(chr), ""))       # blank text is not NA until you convert it

sprintf("%03d", num)                     # put(num, z3.); integer only; NA -> " NA"

as.Date(iso)                             # input(iso, yymmdd10.) as an R Date
as.character(as.Date(iso))               # put(date, yymmdd10.)
format(as.Date(iso), "%Y-%m-%d")         # same character date
as.Date(sas_date, origin = "1960-01-01") # SAS date number -> Date
as.integer(r_date - as.Date("1960-01-01")) # Date -> SAS date number

as.POSIXct(txt, tz = "UTC")              # date only; time after T is dropped
as.POSIXct(txt, format = "%Y-%m-%dT%H:%M:%S", tz = "UTC")
as.POSIXct(sas_dt, origin = "1960-01-01", tz = "UTC")

as.numeric(as.character(fac))            # labels, not factor codes
as.numeric(sub(",", ".", dec_comma, fixed = TRUE)) # one decimal comma, no thousands mark
R SAS
as.character(x) strip(put(x, best.))
as.numeric(x) input(x, best.)
sprintf("%03d", x) put(x, z3.)
as.Date(x) input(x, yymmdd10.)
as.character(d) / format(d, "%Y-%m-%d") put(date, yymmdd10.)
as.Date(n, origin = "1960-01-01") n is already a SAS date
as.POSIXct(n, origin = "1960-01-01", tz = "UTC") n is a SAS datetime (E8601DT.)
as.numeric(sub(",", ".", x, fixed = TRUE)) input(x, commax.) for a single decimal comma

Traps: as.numeric("001") is 1. Leading zeros do not survive a numeric. "" and " " are not NA; as.numeric() of either is NA and does not warn. as.numeric("1,25") is NA and warns NAs introduced by coercion. as.numeric() does not read a decimal comma, and OutDec only changes printing. as.Date(n) with no origin counts from 1970-01-01, not from the SAS epoch 1960-01-01. as.Date() reads; put(date, yymmdd10.) is as.character(d) or format(d, "%Y-%m-%d"). as.numeric() on a factor returns the level code, not the label. put(x, best.) is BEST12. and can contain blanks; as.character() does not pad. sprintf("%03d", x) needs an integer. NA becomes " NA". A non-integer numeric errors. as.POSIXct("2025-06-03T14:30:00", tz = "UTC") with no format does not error: the date parses and the time is dropped (2025-06-03 UTC). Pass format to keep the clock time.

Worked examples below.

Character and numeric

library(dplyr)

Attaching package: 'dplyr'
The following objects are masked from 'package:stats':

    filter, lag
The following objects are masked from 'package:base':

    intersect, setdiff, setequal, union
raw <- data.frame(
  USUBJID = c("STD-001-0001", "STD-001-0002", "STD-001-0003", "STD-001-0004"),
  SITEC = c("001", "010", "002", "003"),
  AGEC = c("54", "61", "", " "),
  RESC = c("1.25", "1,25", "12", "8")
)

raw
       USUBJID SITEC AGEC RESC
1 STD-001-0001   001   54 1.25
2 STD-001-0002   010   61 1,25
3 STD-001-0003   002        12
4 STD-001-0004   003         8

SITEC is character, including leading zeros. AGEC has a zero-length string and a single space. RESC has one decimal point and one decimal comma.

raw |>
  mutate(
    SITE = as.numeric(SITEC),
    AGE = as.numeric(na_if(trimws(AGEC), ""))
  )
       USUBJID SITEC AGEC RESC SITE AGE
1 STD-001-0001   001   54 1.25    1  54
2 STD-001-0002   010   61 1,25   10  61
3 STD-001-0003   002        12    2  NA
4 STD-001-0004   003         8    3  NA

SAS:

site = input(sitec, best.);
age = input(agec, best.);

SITE is 1, 10, 2, 3. "001" and "010" lose their leading zeros because a numeric has no zeros to keep. AGE is 54, 61, then NA for both the empty string and the space. trimws() makes an all-blank string "", and na_if() makes that NA before as.numeric(), so this path does not warn.

Passed straight to as.numeric(), those blanks still become NA. There is no warning. A blank is not the same thing as text that cannot be a number.

as.numeric(c("54", "", " "))
[1] 54 NA NA
as.character(1)
[1] "1"

as.character(1) is "1". SAS strip(put(1, best.)) is "1". put(1, best.) without strip is BEST12., so the result can contain leading blanks. as.character() does not pad.

ImportantBlank is not NA

SAS character missing is blank. missing() is true when a character value is all blanks. In R, "" and " " are present. Only NA is missing. as.numeric("") returns NA and does not warn. SAS input("", best.) also returns numeric missing without an invalid-data note. The character value was not NA before the call.

Decimal comma

raw |>
  mutate(RES = as.numeric(RESC))
Warning: There was 1 warning in `mutate()`.
ℹ In argument: `RES = as.numeric(RESC)`.
Caused by warning:
! NAs introduced by coercion
       USUBJID SITEC AGEC RESC   RES
1 STD-001-0001   001   54 1.25  1.25
2 STD-001-0002   010   61 1,25    NA
3 STD-001-0003   002        12 12.00
4 STD-001-0004   003         8  8.00

1.25, NA, 12, 8. "1,25" is not a number to as.numeric(). mutate() warns NAs introduced by coercion for that value. A blank does not warn; "1,25" does. SAS input("1,25", best.) also returns missing, and it writes an invalid-data note. input("1,25", ?? best.) returns missing and suppresses the note. Neither reads the comma as a decimal mark.

SAS input("1,25", commax5.) reads 1.25. COMMAX uses a comma as the decimal mark and a period as the thousands separator. COMMA is the opposite.

as.numeric(sub(",", ".", "1,25", fixed = TRUE))
[1] 1.25

sub() here replaces one comma. Use it only when the comma is the decimal mark and there is no thousands separator. It is not a general COMMAX parser. as.numeric() does not follow options(OutDec = ","). That option changes how numbers are printed, not how as.numeric() reads them.

Put the zeros back

The numeric value cannot store the width. Write it on the way out, as Zw. does.

raw |>
  mutate(
    SITE = as.integer(SITEC),
    SITE_CHAR = as.character(SITE),
    SITEC2 = sprintf("%03d", SITE)
  )
       USUBJID SITEC AGEC RESC SITE SITE_CHAR SITEC2
1 STD-001-0001   001   54 1.25    1         1    001
2 STD-001-0002   010   61 1,25   10        10    010
3 STD-001-0003   002        12    2         2    002
4 STD-001-0004   003         8    3         3    003

as.integer("001") is 1. as.character() of that value is "1". sprintf("%03d", SITE) is "001". SAS put(site, z3.) writes the same character zeros.

%03d needs an integer. A non-integer numeric errors. Missing becomes " NA", not a blank:

sprintf("%03d", c(1L, NA_integer_))
[1] "001" " NA"

Factors

as.numeric() on a factor uses the level codes.

fac <- factor(c("10", "20"))
as.numeric(fac)
[1] 1 2
as.numeric(as.character(fac))
[1] 10 20

The codes are 1 and 2. The labels are "10" and "20". Convert the labels: as.numeric(as.character(fac)).

Dates

A complete ISO date string and a SAS date number are different types. Partial --DTC strings are a separate problem: see Partial dates. Study day on Date values is Study day.

as.Date(c("2025-06-03", "2025-06-10", ""))
[1] "2025-06-03" "2025-06-10" NA          

as.Date("2025-06-03") is that calendar day. as.Date("") is NA. SAS input("2025-06-03", yymmdd10.) does not return a Date. It returns the integer day count from 1960-01-01. as.Date() is that INPUT direction only.

as.character(as.Date("2025-06-03"))
[1] "2025-06-03"
format(as.Date("2025-06-03"), "%Y-%m-%d")
[1] "2025-06-03"

Both are the character 2025-06-03. That is put(date, yymmdd10.).

sas_n <- as.integer(as.Date("2025-06-03") - as.Date("1960-01-01"))
sas_n
[1] 23895
as.Date(sas_n, origin = "1960-01-01")
[1] "2025-06-03"
as.Date(sas_n)
[1] "2035-06-04"

sas_n is 23895. origin = "1960-01-01" returns 2025-06-03. With no origin, as.Date() counts from 1970-01-01 and returns 2035-06-04.

ImportantA SAS date number is not an R Date

as.Date(n) with no origin means days since 1970-01-01. A SAS date is days since 1960-01-01. The same integer is not the same calendar day. Pass origin = "1960-01-01" when n came from SAS.

The other direction, from an R Date to a SAS date number:

as.integer(as.Date("2025-06-03") - as.Date("1960-01-01"))
[1] 23895

Datetime

SAS datetime values are seconds since 1960-01-01 00:00:00. input("2025-06-03T14:30:00", e8601dt.) reads that ISO form.

txt <- "2025-06-03T14:30:00"
as.POSIXct(txt, tz = "UTC")
[1] "2025-06-03 UTC"
as.POSIXct(txt, format = "%Y-%m-%dT%H:%M:%S", tz = "UTC")
[1] "2025-06-03 14:30:00 UTC"
sas_dt <- as.numeric(
  as.POSIXct("2025-06-03 14:30:00", tz = "UTC") -
    as.POSIXct("1960-01-01", tz = "UTC"),
  units = "secs"
)
as.POSIXct(sas_dt, origin = "1960-01-01", tz = "UTC")
[1] "2025-06-03 14:30:00 UTC"

With no format, as.POSIXct(txt, tz = "UTC") does not error. It returns 2025-06-03 UTC. The date parses and the time is dropped. Pass format = "%Y-%m-%dT%H:%M:%S" to keep 14:30:00.

Set tz. A POSIXct value is an instant. Another zone prints another clock time. "UTC" keeps the clock time stored in a zone-naive SAS datetime.